Getting Started
Start synthesising speech with the Sonex Labs API in under 5 minutes.
Go to sonexlabs.com and sign up. Once logged in, you will be placed in a workspace (tenant).
In the sidebar, go to Configurations → Developer. Under the API Keys tab, click Create key, give it a name, and copy the key immediately; it is shown only once.
vsk_ and are scoped to your workspace. Store them in environment variables, never in source code.Replace YOUR_API_KEY and REPLACE_WITH_VOICE_ID below. You can find voice IDs by calling GET /v1/voices or browsing the voices in your dashboard. Only text is required. voice_id, language, speed, output_format, and sample_rate are all optional.
curl --location 'https://api.sonexlabs.com/v1/speech' \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"text": "Your premium of rupees twelve thousand five hundred is due on the fifteenth of September.",
"voice_id": "REPLACE_WITH_VOICE_ID",
"language": "auto",
"speed": 1.0,
"output_format": "wav",
"sample_rate": 24000
}'language defaults to "auto" (auto-detected from the text). A value other than "auto"is accepted but has no effect on pronunciation. speed accepts 0.5–2.0 (default 1.0); values outside that range are clamped, not rejected. output_format accepts wav, mp3, ogg, opus, or mulaw (default wav). opus is Opus in an Ogg container, typically 10-20x smaller than wav for the same speech; mulaw is G.711 mu-law in a WAV container, for telephony bridges. sample_rate accepts 8000, 16000, 22050, 24000, 44100, or 48000 Hz (default 24000); 8000/16000 are the common choices for telephony and ASR pipelines. text is limited to 5,000 characters (~4-5 minutes of audio) on this endpoint, so use /v1/speech/stream below for longer input.
A successful response returns raw audio bytes, with Content-Type matching output_formatand the character count billed in the X-Chars-Billed header. If voice_id isn't recognized, the request still succeeds using the default platform voice. Check the X-Voice-Fallback response header ("true"/"false") and X-Voice-Fallback-Reason to detect this instead of a silent mismatch.
Prefer a GUI? Run every endpoint from our Postman collection instead. It is pre-filled with the same requests shown on this page.
POST /v1/speech/stream takes the exact same request body as /v1/speech, but streams audio back as chunked HTTP sentence-by-sentence instead of waiting for the full file. Use this for longer scripts to reduce perceived latency.
curl --location 'https://api.sonexlabs.com/v1/speech/stream' \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"text": "Your premium of rupees twelve thousand five hundred is due on the fifteenth of September.",
"voice_id": "REPLACE_WITH_VOICE_ID",
"language": "auto",
"speed": 1.0,
"sample_rate": 24000
}'output_format is not accepted here. Streamed audio is always chunked WAV/PCM. sample_rate is honored (same allowed values as /v1/speech). text is limited to 20,000 characters (~15-20 minutes of audio) on this endpoint.
Billing: your full request is pre-authorized against your balance upfront (so you can't start a request you can't afford). Complete the stream and you're billed the full requested amount. Disconnect early and you're billed a bounded estimate of how much audio actually reached you before the drop, never the full request for a stream that only partially delivered. X-Chars-Billed shows the requested/authorized amount; the amount actually charged for an early-disconnected stream will be lower and shows up in your usage ledger.
POST /v1/transcribe takes a finished recording and returns its transcript. Send the audio as the raw body with a matching Content-Type, or as a multipart upload under file. WAV, MP3, FLAC, OGG, Opus, WebM, AAC and M4A are accepted, up to 10 MB and 10 minutes.
curl --location 'https://api.sonexlabs.com/v1/transcribe' \ --header 'Authorization: Bearer YOUR_API_KEY' \ --header 'Content-Type: audio/wav' \ --data-binary '@recording.wav'
The language is detected for you, including recordings that switch between languages, so there is nothing to configure. The reply carries text, the duration_seconds of the audio and a request_id to quote if you ever need to ask about a request.
Billing: per second of audio, for the recording's real duration. Audio we reject before transcribing it, because the format is unsupported or the file is too large, too long or unreadable, is never billed. Send an Idempotency-Key header and a retry within 10 minutes returns the transcript again with seconds_billed of 0.
wss://api.sonexlabs.com/v1/transcribe/stream takes 16 kHz mono 16-bit PCM audio as binary frames and sends back partial transcripts as they form and final ones as each utterance completes. Send {"type":"end"} when the speaker stops.
import websockets, json
async with websockets.connect(
"wss://api.sonexlabs.com/v1/transcribe/stream",
additional_headers={"Authorization": "Bearer YOUR_API_KEY"},
) as ws:
# Wait for {"type": "ready"} before streaming, so the session is open and
# the first words are transcribed rather than queued behind the handshake.
while json.loads(await ws.recv())["type"] != "ready":
pass
for frame in pcm_frames: # 16 kHz mono PCM16, ~100 ms each, in real time
await ws.send(frame)
await ws.send(json.dumps({"type": "end"}))
async for message in ws:
event = json.loads(message)
if event["type"] == "final":
print(event["text"])
elif event["type"] == "done":
print("billed", event["seconds_billed"], "seconds")
breakA browser cannot set headers on a WebSocket, so a key may also be passed as ?api_key=vsk_.... A key sent that way is visible to anything that can read the URL, so mint a short-lived one from your own backend rather than shipping your main key to a page.
Billing: per second for as long as the connection is open, whether or not anyone is speaking, because a streaming transcriber charges for the open connection. The total arrives as seconds_billed on the done message. Close the socket as soon as you are finished. A connection is closed for you after 10 minutes, or after 30 seconds with no recognised speech.
Fetch the voices available to your workspace. The id field is what you pass as voice_id.
curl https://api.sonexlabs.com/v1/voices \ -H "Authorization: Bearer YOUR_API_KEY"
[
{
"id": "72ly9crx9v",
"name": "Alok",
"language": "Hindi",
"languages": ["Hindi"],
"gender": "male",
"provider": "panini",
"type": "platform",
"preview_url": "https://api.sonexlabs.com/v1/voices/72ly9crx9v/preview",
"tags": ["Indian", "Natural, Professional"],
"created_at": null
}
]preview_url plays sample audio for the voice. It needs no API key, so it can be used directly as the src of an <audio> element. Request it rather than storing it: it redirects to storage that can move.
POST /v1/voices/clone is a multipart/form-data request, not JSON. Upload the reference audio directly with the file field (recommended), or pass audio_url if it is already hosted somewhere; provide exactly one of the two. Minimum 10 seconds of clean audio, one speaker. There is no language field: Pāṇini auto-detects language at synthesis time, same as with platform voices.
curl -X POST https://api.sonexlabs.com/v1/voices/clone \ -H "Authorization: Bearer YOUR_API_KEY" \ -F "name=My Custom Voice" \ -F "file=@reference.wav"
Runs synchronously; the response is the ready-to-use voice, no polling required.
{
"id": "v_xyz789",
"name": "My Custom Voice",
"language": null,
"languages": [],
"gender": null,
"provider": "custom",
"type": "custom",
"preview_url": "https://api.sonexlabs.com/v1/voices/v_xyz789/preview",
"tags": ["custom", "cloned"],
"created_at": "2026-08-05T18:19:22.468Z"
}type is "platform" for built-in Sonex Labs voices and "custom" for voices your workspace cloned. Use it to tell the two apart when you list voices with GET /v1/voices.
To remove a cloned voice, call DELETE /v1/voices/{voice_id}:
curl --location --request DELETE 'https://api.sonexlabs.com/v1/voices/REPLACE_WITH_VOICE_ID' --header 'Authorization: Bearer YOUR_API_KEY'
Each synthesis call deducts from your TTS service credits first, then from your wallet balance. Check your remaining balance at any time:
curl https://api.sonexlabs.com/v1/balance \ -H "Authorization: Bearer YOUR_API_KEY"
{
"wallet": { "balance": 5.00, "currency": "USD" },
"credits": [
{ "service": "tts", "available": 50000, "next_expiry": "2027-06-01T00:00:00Z" }
]
}GET /v1/usage summarises a period rather than making you download every call. Costs use the same components as a call and are rounded the same way, so a day in the breakdown equals the sum of that day's calls from /v1/calls. Group by day or by agent, and set timezone to your own or a late evening call will be counted on the following day.
curl 'https://api.sonexlabs.com/v1/usage?from=2026-09-01&to=2026-09-30&timezone=Asia/Kolkata' \ -H "Authorization: Bearer YOUR_API_KEY"
{
"period": { "from": "2026-08-31T18:30:00+00:00", "to": "2026-09-30T18:29:59+00:00", "timezone": "Asia/Kolkata" },
"totals": {
"cost_usd": 0.272111,
"cost": { "llm": 0.000362, "stt": 0.004987, "tts": 0.113552, "telephony": 0.0 },
"calls": {
"total": 53,
"connected": 52,
"duration_secs": 4546,
"by_status": { "completed": 52, "no-answer": 1 },
"by_direction": { "outbound": 48, "inbound": 5 }
},
"speech": { "requests": 0, "characters": 0 }
},
"group_by": "day",
"breakdown": [
{
"date": "2026-09-21",
"cost_usd": 0.219522,
"cost": { "llm": 0.000198, "stt": 0.004027, "tts": 0.062087, "telephony": 0.0 },
"calls": { "total": 27, "connected": 26, "duration_secs": 2378 },
"speech": { "requests": 0, "characters": 0 }
}
]
}Balance is not repeated here. Use GET /v1/balance for what remains.
GET /v1/calls returns the calls your voice agents handled, newest first, scoped to the tenant that owns the API key. Pagination is cursor based: request the first page without cursor, then pass the next_cursor from each response until has_more is false. Filters are available for agent_id, status, direction, phone_number, started_after and started_before.
curl 'https://api.sonexlabs.com/v1/calls?limit=25&direction=outbound' \ -H "Authorization: Bearer YOUR_API_KEY"
{
"data": [
{
"id": "CA7f3a1b90c24e4d02",
"status": "completed",
"direction": "outbound",
"agent": { "id": "3f0c9d18-1c4a-4a3a-9a75-7e1c2b6f42aa", "name": "Support Line" },
"from": "+14155550142",
"to": "+916301979825",
"started_at": "2026-09-20T11:42:07.482Z",
"ended_at": "2026-09-20T11:42:56.117Z",
"duration_secs": 49,
"cost_usd": 0.014226,
"has_recording": true,
"has_transcript": true
}
],
"has_more": true,
"next_cursor": "WyIyMDI2LTA5LTIwVDExOjQyOjA3WiIsIjhkMWEiXQ"
}The id is the provider call SID, so it matches the identifier in your telephony provider's own logs. Pass it to GET /v1/calls/{call_id} for the full record: the models used, the itemised cost, the transcript and every tool the agent invoked. Use include=transcript,tool_calls to choose which sections are returned, or include= for the summary fields only.
curl 'https://api.sonexlabs.com/v1/calls/CA7f3a1b90c24e4d02' \ -H "Authorization: Bearer YOUR_API_KEY"
The transcript is also available on its own at GET /v1/calls/{call_id}/transcript, which is the cheaper choice for a client that polls.
GET /v1/calls/{call_id}/recording returns the audio. By default it responds with 302 Found and a Location header, so it works directly as the src of an audio player. Pass redirect=false to receive the link as JSON instead. Links expire after 15 minutes, so store the call_id and request a fresh link when you need one rather than caching the URL.
curl 'https://api.sonexlabs.com/v1/calls/CA7f3a1b90c24e4d02/recording?redirect=false' \ -H "Authorization: Bearer YOUR_API_KEY"
{
"id": "CA7f3a1b90c24e4d02",
"url": "https://storage.sonexlabs.com/storage/v1/object/sign/call-recordings/...",
"expires_at": "2026-09-20T12:12:56.117Z",
"duration_secs": 49
}Pāṇini TTS supports 250+ languages and always auto-detects the language from your input text. No configuration needed. The optional language field on POST /v1/speech and POST /v1/speech/stream defaults to "auto"; leave it as "auto" to auto-detect.
Full list of 250+ languages at sonexlabs.com/models/tts.
The API enforces two limits per tenant, shared across all API keys in your workspace: 5 requests/second (burst) and 120 requests/minute (sustained). When exceeded, the API returns 429 Too Many Requests. Retry after the number of seconds indicated in the Retry-After response header.
Two client-side changes remove latency that has nothing to do with synthesis itself:
A fresh connection pays DNS lookup + TCP connect + TLS handshake before your request is even sent. measured at ~150-300ms combined. Reuse one HTTP client/session across requests (e.g. requests.Session() in Python, http.Agent({ keepAlive: true }) in Node, or any pooled aiohttp.ClientSession) instead of opening a new connection per call. This cost is paid once per connection, not once per request.
output_format for slower linksoutput_format: "wav" is the default and is uncompressed, so the response payload is larger and takes longer to download over higher-latency or bandwidth-constrained connections (e.g. mobile networks). If you don't need raw WAV, request "opus" instead: same synthesis time on our side, typically 10-20x smaller payload, faster download. "mp3" and "ogg"are also available if your playback stack needs them specifically, but compress less than "opus".