Speech and audio AI.
Whisper-large fits in 8GB VRAM. Real-time TTS and music models fit on consumer GPUs. Training novel models needs workstation-class hardware.
Top GPUs for this workload.
Ranked by suitability — higher fitness scores mean the card handles this workload more comfortably.
Top models for this workload.
Whisper Large v3
OpenAI's open-weight speech-to-text — the standard transcription model.
Whisper Medium
769M Whisper variant — half the size of Large, 80% of the accuracy.
Whisper Small
244M Whisper — fits on edge GPUs and CPU.
GPT-4o
OpenAI's multimodal model — text, vision, audio in one.
Whisper Base
74M Whisper — browser / Raspberry Pi-deployable.
Whisper Tiny
39M Whisper — runs in-browser via WebGPU.
Claude Haiku 4.5
Fast, cheap Claude variant for high-throughput inference.