provider · Deepgram
Deepgram (Nova-3)
Deepgram's Nova-3 model is widely regarded as the most accurate proprietary speech-to-text engine available: it has the lowest word-error-rate among commercial transcription APIs and handles both pre-recorded audio (batch) and live microphone streams (real-time) well. For a student this is the tool behind transcribing interviews, lab meeting recordings, or a live captioning feature in an app — anywhere accuracy on technical vocabulary and noisy audio matters more than cost. It's also cheap at scale: batch transcription runs about $0.0043 per minute of audio, and streaming about $0.0077 per minute, which is inexpensive compared to hiring a human transcriber but still an ongoing cost compared to running something locally. Like the other developer-facing tools in this chapter, using it means calling an API rather than clicking a button on a website.
What it can help with
- speech-to-text
- transcription-api
- batch-transcription
- real-time-captioning
- technical-vocabulary
- noisy-audio
- cloud-api
- low-cost
- streaming-transcription
- high-accuracy