Skip to main content
47 STT model(s) from 7 provider(s). Priced from input_audio_kseconds, US dollars per 1,000 seconds of input audio. Shown per minute, the unit vendors quote. Pick a vendor in the sidebar to see that vendor alone. Streaming and batch ship as separate rows, as do language tiers, because vendors price them separately. 2 further STT model(s) bill audio as tokens rather than by the second and are listed on their vendor page: gpt-4o-mini-transcribe, gpt-4o-transcribe.

Via LiveKit Inference

What the same model costs going direct to the vendor versus through the gateway. Scale is the discounted tier. 8 further LiveKit STT model(s) have no direct-vendor entry to compare against. What verified, imported, and seed mean is on How Fresh?.