🎙️ Whisper models
ModelVRAMSpeedAccuracyRecommended
tiny~1 GB⚡⚡⚡⚡★★No
base~1 GB⚡⚡⚡★★★No
small~2 GB⚡⚡★★★★Quick demo
medium~5 GB⚡★★★★Good trade-off
large-v3~10 GB★★★★★✅ RTX 5090
💾 ZFS Dataset
DatasetContent
whisperWhisper models (checkpoints)
🐍 Python example
python
import whisper

# Charger le modèle (téléchargé si absent)
model = whisper.load_model('large-v3')

# Transcrire un fichier audio
result = model.transcribe(
    'enregistrement.mp3',
    language='fr',
    verbose=True
)

print(result['text'])

# Avec timestamps
for segment in result['segments']:
    print(f"[{segment['start']:.1f}s] {segment['text']}")
🎯 Skills in progress
  • Whisper installation with GPU acceleration (CUDA)
  • Audio/video transcription in French and multilingual
  • Model selection based on VRAM/speed constraints
  • Integration into an n8n pipeline (audio trigger → STT → LLM)
  • Multimodal AI: audio → text → LLM → answer

Related pages

References & Sources

CategoryResourceURL
Official documentationOpenAI Whispergithub.com/openai/whisper
ModelsWhisper — modèles multilingues (Hugging Face)huggingface.co/openai
StatusTranscription audio/vidéo — validé, portage planifié sur VM210—
LicenseWhisper — MITgithub.com/openai/whisper — LICENSE
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0