Audio

Create audio transcription request

POST
/audio/transcriptions

Transcribes audio into text

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Request Body

multipart/form-data

file*|

Audio file upload or public HTTP/HTTPS URL. Supported formats: .wav, .mp3, .m4a, .webm, .flac, .ogg, .opus, .aac. Maximum duration 4 hours; longer audio is rejected with audio_too_long. Binary uploads are additionally capped at 80 MB (HTTP 413); URL-fetched audio is capped at 1 GB.

model?"openai/whisper-large-v3"

Model to use for transcription

Default"openai/whisper-large-v3"

Value in

  • "openai/whisper-large-v3"
language?string

Optional ISO 639-1 language code. If auto is provided, language is auto-detected.

Default"en"
prompt?string

Optional text to bias decoding. Supported only on Whisper-family models (e.g. openai/whisper-large-v3). Other STT models (e.g. nvidia/parakeet-tdt-0.6b-v3) accept the field for API compatibility but ignore it.

response_format?string

The format of the response

Default"json"

Value in

  • "json"
  • "verbose_json"
temperature?number

Sampling temperature between 0.0 and 1.0

Range0 <= value <= 1
Default0
timestamp_granularities?|array<>

Controls level of timestamp detail in verbose_json. Only used when response_format is verbose_json. Can be a single granularity or an array to get multiple levels.

Default"segment"
diarize?boolean

Whether to enable speaker diarization. When enabled, you will get the speaker id for each word in the transcription. In the response, in the words array, you will get the speaker id for each word. In addition, we also return the speaker_segments array which contains the speaker id for each speaker segment along with the start and end time of the segment along with all the words in the segment. For eg - ... "speaker_segments": [ "speaker_id": "SPEAKER_00", "start": 0, "end": 30.02, "words": [ { "id": 0, "word": "Tijana", "start": 0, "end": 11.475, "speaker_id": "SPEAKER_00" }, ...

Defaultfalse
min_speakers?integer

Minimum number of speakers expected in the audio. Used to improve diarization accuracy when the approximate number of speakers is known.

max_speakers?integer

Maximum number of speakers expected in the audio. Used to improve diarization accuracy when the approximate number of speakers is known.

Response Body

application/json

application/json

application/json

text/html

application/json

curl -X POST "https://example.com/audio/transcriptions" \  -F file="string"
{  "text": "string"}

Real-time text-to-speech via WebSocket GET

Establishes a WebSocket connection for real-time text-to-speech generation. This endpoint uses WebSocket protocol (wss://api.together.ai/v1/audio/speech/websocket) for bidirectional streaming communication. **Connection Setup:** - Protocol: WebSocket (wss://) - Authentication: Pass API key as Bearer token in Authorization header - Parameters: Sent as query parameters (model, voice, max_partial_length, language) **Client Events:** - `tts_session.updated`: Update session parameters like voice. The `session` object also accepts an `extra_params` field for additional model-specific parameters that fine-tune speech generation behavior, such as `pronunciation_dict` (a list of pronunciation rules for specific characters or symbols, where each entry uses the format `"<source>/<replacement>"` (e.g., `["omg/oh my god"]`) to override how the model pronounces matching tokens). ```json { "type": "tts_session.updated", "session": { "voice": "tara", "extra_params": { "pronunciation_dict": ["omg/oh my god"] } } } ``` - `input_text_buffer.append`: Send text chunks for TTS generation ```json { "type": "input_text_buffer.append", "text": "Hello, this is a test." } ``` - `input_text_buffer.clear`: Clear the buffered text ```json { "type": "input_text_buffer.clear" } ``` - `input_text_buffer.commit`: Signal end of text input and process remaining text ```json { "type": "input_text_buffer.commit" } ``` **Server Events:** - `session.created`: Initial session confirmation (sent first) ```json { "event_id": "evt_123456", "type": "session.created", "session": { "id": "session-id", "object": "realtime.tts.session", "modalities": ["text", "audio"], "model": "hexgrad/Kokoro-82M", "voice": "tara" } } ``` - `conversation.item.input_text.received`: Acknowledgment that text was received ```json { "type": "conversation.item.input_text.received", "text": "Hello, this is a test." } ``` - `conversation.item.audio_output.delta`: Audio chunks as base64-encoded data ```json { "type": "conversation.item.audio_output.delta", "item_id": "tts_1", "delta": "<base64_encoded_audio_chunk>" } ``` - `conversation.item.audio_output.done`: Audio generation complete for an item ```json { "type": "conversation.item.audio_output.done", "item_id": "tts_1" } ``` - `conversation.item.tts.failed`: Error occurred ```json { "type": "conversation.item.tts.failed", "error": { "message": "Error description", "type": "invalid_request_error", "param": null, "code": "invalid_api_key" } } ``` **Text Processing:** - Partial text (no sentence ending) is held in buffer until: - We believe that the text is complete enough to be processed for TTS generation - The partial text exceeds `max_partial_length` characters (default: 250) - The `input_text_buffer.commit` event is received **Audio Format:** - Format: Raw PCM (s16le, mono) - Sample Rate: 24000 Hz - Encoding: Base64 (per delta event) - Delivered via `conversation.item.audio_output.delta` events **Error Codes:** - `invalid_api_key`: Invalid API key provided (401) - `missing_api_key`: Authorization header missing (401) - `model_not_available`: Invalid or unavailable model (400) - Invalid text format errors (400)

Create audio translation request POST

Translates audio into English