Skip to main content

Direct answer

Audio transcription and speech generation are separate tasks and can use different models and endpoints. Confirm transcription, speech, audio, or realtime models visible to the current account in the catalog, then follow the model-specific page or console endpoint. This page no longer describes whisper-1, TTS, or one endpoint as universally verified for production. This page was last verified on September 2, 2026.

Common task entry points

Minimal transcription request

Use this only when the console explicitly shows that the model uses the transcription endpoint:

Acceptance items

  • supported audio format, sample rate, channels, size, and duration;
  • language detection, timestamps, segments, and output format;
  • actual audio MIME type and file content;
  • errors for oversized, corrupt, silent, and unsupported files;
  • whether retrying the same file can create duplicate charges;
  • usage, billing unit, and call logs.
A catalog entry does not prove that every audio endpoint works. Official upstream support also does not prove that the current gateway group has integrated it. Acceptance-test the target API key and representative audio before production.