🎙️ Whisper models
| Model | VRAM | Speed | Accuracy | Recommended |
|---|---|---|---|---|
| tiny | ~1 GB | ⚡⚡⚡⚡ | ★★ | No |
| base | ~1 GB | ⚡⚡⚡ | ★★★ | No |
| small | ~2 GB | ⚡⚡ | ★★★★ | Quick demo |
| medium | ~5 GB | ⚡ | ★★★★ | Good trade-off |
| large-v3 | ~10 GB | ★★★★★ | ✅ RTX 5090 |
💾 ZFS Dataset
| Dataset | Content |
|---|---|
| whisper | Whisper models (checkpoints) |
🐍 Python example
python
import whisper
# Charger le modèle (téléchargé si absent)
model = whisper.load_model('large-v3')
# Transcrire un fichier audio
result = model.transcribe(
'enregistrement.mp3',
language='fr',
verbose=True
)
print(result['text'])
# Avec timestamps
for segment in result['segments']:
print(f"[{segment['start']:.1f}s] {segment['text']}") 🎯 Skills in progress
- Whisper installation with GPU acceleration (CUDA)
- Audio/video transcription in French and multilingual
- Model selection based on VRAM/speed constraints
- Integration into an n8n pipeline (audio trigger → STT → LLM)
- Multimodal AI: audio → text → LLM → answer