> ## Documentation Index
> Fetch the complete documentation index at: https://docs.60db.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# STT WebSocket

> Real-time Speech-to-Text WebSocket API for streaming audio transcription

# STT WebSocket API

Real-time Speech-to-Text transcription via WebSocket streaming with support for 39 languages (including code-switched Indic+English) and telephony integration. Powered by 60db STT v01 (a non-hallucinating, multi-backend speech recognition stack).

## 🚀 Quick Start (Copy & Paste)

```javascript theme={null}
const WebSocket = require('ws');

// 1. Your API key
const API_KEY = 'sk_live_your_api_key';

// 2. Connect
const ws = new WebSocket(`wss://api.60db.ai/ws/stt?apiKey=${API_KEY}`);

// 3. Handle messages
ws.on('message', (data) => {
  const msg = JSON.parse(data);

  // Authenticated? Start session!
  if (msg.connection_established) {
    console.log('✅ Authenticated');
    ws.send(JSON.stringify({
      type: 'start',
      languages: ['en'],
      config: { encoding: 'mulaw', sample_rate: 8000, continuous_mode: true }
    }));
  }

  // Session ready? Send audio!
  if (msg.type === 'connected') {
    console.log('✅ Ready! Send audio now');

    // Send dummy audio (480 bytes every 60ms)
    let count = 0;
    const interval = setInterval(() => {
      ws.send(Buffer.alloc(480, 0xff));
      if (++count >= 83) {  // 5 seconds
        clearInterval(interval);
        ws.send(JSON.stringify({ type: 'stop' }));
      }
    }, 60);
  }

  // Got text!
  if (msg.type === 'transcription' && msg.is_final) {
    console.log('📝', msg.text);
  }

  // Done!
  if (msg.type === 'session_stopped') {
    console.log('✅ Complete! Cost:', msg.billing_summary.total_cost);
    ws.close();
  }
});
```

**That's it!** You'll see:

* ✅ Authenticated
* ✅ Ready! Send audio now
* 📝 Hello world (transcribed text)
* ✅ Complete! Cost: \$0.000043

***

## 📖 How It Works (5 Simple Steps)

1. **Connect** with your API key
2. **Send** `{ type: "start", ... }` to begin session
3. **Stream** audio data (binary chunks)
4. **Receive** text transcriptions in real-time
5. **Stop** with `{ type: "stop" }` when done

***

## Endpoint

<ParamField name="url" type="string">
  `ws://api.60db.ai/ws/stt` or `wss://api.60db.ai/ws/stt`
</ParamField>

## Authentication

Query parameter authentication:

<ParamField name="apiKey" type="string" required>
  Your API key for authentication
</ParamField>

Example:

```
ws://api.60db.ai/ws/stt?apiKey=sk_live_your_api_key
```

## Connection Details

| Property       | Value                                                |
| -------------- | ---------------------------------------------------- |
| Protocol       | WebSocket (RFC 6455)                                 |
| Frame types    | Binary (telephony) or Text/JSON (browser)            |
| Ping/keepalive | Server sends WebSocket pings every 30s (timeout 10s) |

## Session Lifecycle

```
Client                          Server
  │                               │
  │──── TCP/TLS connect ─────────►│
  │◄─── {"connecting": true} ─────│  authenticating...
  │◄─── connection_established ──│  authenticated
  │                               │
  │──── {"type":"start", ...} ───►│
  │◄─── {"type":"connected"} ─────│  session ready
  │                               │
  │──── audio frames / messages ─►│
  │◄─── {"type":"speech_started"} │  VAD detected voice
  │◄─── {"type":"transcription"}  │  is_final=true, speech_final=false (first emit, context only)
  │◄─── {"type":"transcription"}  │  is_final=true, speech_final=true  (canonical answer)
  │                               │
  │──── {"type":"stop"} ─────────►│
  │◄─── {"type":"session_stopped"}│
```

**Two-phase finals (context-gated LLM refinement).** When you supply a `context` object on `start`, every utterance produces **two** `transcription` events sharing a `sentence_id`:

1. **First emit** — `is_final: true, speech_final: false` — fast dict-corrected text. Use for low-latency UI paint and barge-in.
2. **Canonical** — `is_final: true, speech_final: true` — definitive LLM-refined answer. Always arrives.

When `context` is **omitted**, every utterance produces a single `transcription` event with `is_final: true, speech_final: true` (no first emit). Simple consumers can gate exclusively on `speech_final: true` and ignore the rest — that gives them exactly one canonical event per utterance regardless of whether refinement is on.

## Client → Server Messages

### `start` — Begin session

Sent once after connection is established. Must be sent before any audio.

<RequestExample>
  ```json Request theme={null}
  {
    "type": "start",
    "languages": ["en", "hi"],
    "context": {
      "general": [
        { "key": "domain",  "value": "Healthcare" },
        { "key": "doctor",  "value": "Dr. Martha Smith" }
      ],
      "text":  "Routine diabetes follow-up consultation.",
      "terms": ["Celebrex", "Zyrtec", "Metformin", "HbA1c"]
    },
    "config": {
      "encoding": "mulaw",
      "sample_rate": 8000,
      "utterance_end_ms": 500,
      "continuous_mode": true,
      "interim_results_frequency": 300,
      "audio_enhancement": "adaptive",
      "diarize": false,
      "remove_fillers": false
    }
  }
  ```
</RequestExample>

**Parameters:**

<ParamField name="languages" type="array">
  Array of ISO 639-1 codes from the supported set (see `GET /stt/languages`), e.g. `["en", "hi"]`. Max 5 per session. Omit or send `null` to auto-detect across all 39 v1 languages. Arabic dialect tags (`ar-eg`, …) are rejected. Unsupported: `ur`, `ja`, `ko`, `zh`, `th`, `vi`, `id`, `tl`, `sw`, `tr`, `fa`, `he`.

  **Important:** never send `languages: "auto"` or `["auto"]`. The auto entry in `GET /stt/languages` is a convenience for the REST `/stt` form-upload flow; on WebSocket you must use `null` instead. Sending the literal string `"auto"` returns `language 'auto' is not in the v1 supported`. The 60db `/ws/stt` proxy strips it automatically as a safety net, but your client should send `null` directly.
</ParamField>

<ParamField name="context" type="object">
  Optional hint object `{general, text, terms}` that opens the server-side LLM refinement gate. When supplied, each utterance emits **two** `transcription` events sharing a `sentence_id` — a fast first emit (`speech_final: false`) with dict-corrected text, followed \~300–700 ms later by a canonical emit (`speech_final: true`) with LLM-refined text (proper nouns corrected, fillers removed, punctuation added, script consistency enforced). See [Canonical-answer semantics](#canonical-answer-semantics-speech-final) below.

  * `general` — array of `{key, value}` pairs. Free-form metadata (`domain`, `topic`, speaker names) surfaced to the LLM verbatim as hint lines.
  * `text` — background paragraph describing the session. Useful for narrative context.
  * `terms` — array of proper nouns, acronyms, and domain-specific jargon to preserve in the transcript.

  All three fields are optional; at least one should be populated. Omit `context` entirely to disable refinement for the session — each utterance then arrives as a single `speech_final: true` event (no first emit).
</ParamField>

<ParamField name="config.encoding" type="string" default="mulaw">
  Audio encoding format. Use `"mulaw"` for telephony/Twilio, `"linear"` for browser PCM. Options: `"mulaw"`, `"linear"`
</ParamField>

<ParamField name="config.sample_rate" type="integer" default="8000">
  Actual sample rate of the audio being sent. Server resamples to 16kHz internally. Options: `8000`, `16000`, `24000`, `44100`, `48000`
</ParamField>

<ParamField name="config.utterance_end_ms" type="integer" default="500">
  Silence duration (ms) after last speech chunk before finalizing the utterance. **Minimum 300ms**. Recommended 500-1000ms for voicebots. Range: `≥ 300`
</ParamField>

<ParamField name="config.continuous_mode" type="boolean" default="false">
  Keep session alive between utterances instead of stopping after first transcription. Required for voicebot / phone call use cases.
</ParamField>

<ParamField name="config.interim_results_frequency" type="integer">
  How often (ms) to emit interim (partial) transcription results during speech. Use 300ms for barge-in, 500ms otherwise. Disabled by default. Range: `≥ 300`
</ParamField>

<ParamField name="config.diarize" type="boolean" default="false">
  Run pyannote speaker diarization on each finalized utterance and attach a `speakers` array. Requires `HF_TOKEN` on the server. Adds \~50–150 ms latency per final.
</ParamField>

<ParamField name="config.min_speakers" type="integer | null">
  Lower bound on diarization speaker count. Only read when `diarize=true`.
</ParamField>

<ParamField name="config.max_speakers" type="integer | null">
  Upper bound on diarization speaker count. Only read when `diarize=true`.
</ParamField>

<ParamField name="config.audio_enhancement" type="string" default="off">
  Real-time audio enhancement to improve transcription quality on noisy input. There are **3 options** available:

  | Value        | Description                                                                                        |
  | ------------ | -------------------------------------------------------------------------------------------------- |
  | `"off"`      | No audio processing (default)                                                                      |
  | `"light"`    | Noise reduction only — best for mildly noisy environments                                          |
  | `"adaptive"` | Automatic noise reduction + gain control based on input levels — best for variable/telephony audio |

  Use `"adaptive"` for telephony or noisy environments, `"light"` when you only need mild cleanup, and `"off"` when the input is already clean.
</ParamField>

<ParamField name="config.remove_fillers" type="boolean" default="false">
  Ask the LLM refinement pass to strip filler words (`um`, `uh`, `like`, `you know`, …) from the canonical transcript. Only takes effect when `context` is set — refinement is gated on context, and the raw first-emit still contains the fillers. Non-boolean values are coerced to `false` by the 60db proxy.
</ParamField>

<ParamField name="config.no_speech_threshold" type="float" default="0.60">
  Reserved for legacy client compatibility. **Ignored by the 60db STT backend** — the non-hallucinating backends don't emit a `no_speech_prob`.
</ParamField>

### `audio` — JSON audio chunk (browser mode)

<RequestExample>
  ```json Request theme={null}
  {
    "type": "audio",
    "audio": "<base64-encoded Int16 PCM or μ-law bytes>",
    "encoding": "linear",
    "sample_rate": 48000,
    "timestamp": 1700000000000
  }
  ```
</RequestExample>

**Fields:**

<ParamField name="type" type="string" required>
  Must be `"audio"`
</ParamField>

<ParamField name="audio" type="string" required>
  Base64-encoded audio bytes (Int16 PCM or μ-law)
</ParamField>

<ParamField name="encoding" type="string" required>
  `"linear"` or `"mulaw"`
</ParamField>

<ParamField name="sample_rate" type="integer" required>
  Actual sample rate of the audio
</ParamField>

<ParamField name="timestamp" type="integer">
  Unix ms timestamp — useful for latency measurement
</ParamField>

### Binary frame — raw μ-law audio (telephony mode)

Send a raw WebSocket binary frame with μ-law bytes, no JSON wrapper.
The server auto-detects this as telephony mode on the first binary frame.

```
Recommended chunk size: 480 bytes = 60ms at 8kHz
Twilio default: 160 bytes = 20ms — batch 3 chunks into 60ms before sending
```

### `config` — Change language mid-session

<RequestExample>
  ```json Request theme={null}
  {
    "type": "config",
    "languages": ["hi"],
    "continuous_mode": true
  }
  ```
</RequestExample>

Both `languages` and `continuous_mode` are optional; include only fields you want to change.
Send `"languages": null` to revert to auto-detect.

### `stop` — End session

<RequestExample>
  ```json Request theme={null}
  {
    "type": "stop"
  }
  ```
</RequestExample>

Server processes any remaining audio buffer, sends `session_stopped`, then closes.

### `test` — Ping / latency check

<RequestExample>
  ```json Request theme={null}
  {
    "type": "test",
    "message": "ping",
    "timestamp": 1700000000000
  }
  ```
</RequestExample>

Server echoes `test_response` with the same `timestamp` for round-trip measurement.

## Server → Client Messages

### `connecting` — Authentication in progress

<ResponseExample>
  ```json Response theme={null}
  {
    "connecting": true,
    "message": "Authenticating...",
    "timestamp": 1775465918269
  }
  ```
</ResponseExample>

### `connection_established` — Authentication successful

<ResponseExample>
  ```json Response theme={null}
  {
    "connection_established": {
      "service": "stt",
      "user_id": 43,
      "credit_balance": 9.97,
      "workspace": "default"
    }
  }
  ```
</ResponseExample>

**Fields:**

<ResponseField name="service" type="string">
  Service name: `"stt"`
</ResponseField>

<ResponseField name="user_id" type="integer">
  Your user ID
</ResponseField>

<ResponseField name="credit_balance" type="number">
  Available credits
</ResponseField>

<ResponseField name="workspace" type="string">
  Workspace name
</ResponseField>

### `connected` — After `start` message is processed

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "connected",
    "server_info": {
      "server_type": "60db STT",
      "device": "cuda",
      "model": "60db-stt-v01",
      "processing_mode": "sentence_based_modular",
      "supported_languages": { "en": "English", "hi": "Hindi" },
      "total_languages": 40,
      "features": {
        "sentence_based_processing": true,
        "real_time_streaming": true,
        "telephony_support": true,
        "unicode_support": true,
        "mixed_language_support": true
      }
    }
  }
  ```
</ResponseExample>

### `speech_started` — VAD detected voice activity

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "speech_started",
    "timestamp": 1700000000.123
  }
  ```
</ResponseExample>

Use this for barge-in: interrupt TTS playback when this arrives.
Fired after 2 consecutive VAD-positive chunks (\~64ms of confirmed speech).

### `transcription` — Transcription result

All results (interim and final) share the same `transcription` type — differentiate with flags.

**Final result** (`is_final=true`, `speech_final=true`):

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "transcription",
    "text": "Hello, how are you?",
    "confidence": 0.87,
    "language": "en",
    "language_name": "EN",
    "is_final": true,
    "speech_final": true,
    "is_partial": false,
    "sentence_id": 3,
    "processing_mode": "sentence_complete",
    "duration": 1.82,
    "latency": 0.43,
    "timestamp": 1700000000.456,
    "words": [
      { "word": "Hello", "start": 0.0, "end": 0.32, "confidence": 0.94 },
      { "word": "how",   "start": 0.35, "end": 0.52, "confidence": 0.92 }
    ],
    "utterance_end_ms": 1820
  }
  ```
</ResponseExample>

**Empty speech\_final signal** (`text=""`, `is_final=true`, `speech_final=true`):
Sent when audio was detected but transcription was rejected (silence, hallucination, low confidence, wrong language).
Client should reset its state on this message and not treat it as an error.

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "transcription",
    "text": "",
    "confidence": 0.0,
    "is_final": true,
    "speech_final": true,
    "processing_mode": "speech_end_no_result",
    "timestamp": 1700000000.789
  }
  ```
</ResponseExample>

**Interim result** (`is_final=false`, `speech_final=false`) — only sent when `interim_results_frequency` is set:

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "transcription",
    "text": "Hello how",
    "confidence": 0.72,
    "language": "en",
    "is_final": false,
    "speech_final": false,
    "is_partial": true
  }
  ```
</ResponseExample>

Use interims only for barge-in word-count checks. Never send interim text to the LLM — a final with `is_final=true, speech_final=true` will follow.

**Response Fields:**

<ResponseField name="text" type="string">
  Transcribed text. Empty string = speech-end-no-result signal.
</ResponseField>

<ResponseField name="confidence" type="number">
  0.0–1.0. Telephony typically 0.35–0.75; browser 0.55–0.95.
</ResponseField>

<ResponseField name="language" type="string">
  Detected language code e.g. `"en"`.
</ResponseField>

<ResponseField name="language_name" type="string">
  Uppercase language code e.g. `"EN"`.
</ResponseField>

<ResponseField name="is_final" type="boolean">
  `true` = end of speech reached. May still be followed by a canonical upgrade if LLM refinement is active.
</ResponseField>

<ResponseField name="speech_final" type="boolean">
  `true` = canonical answer, will not be revised. When LLM refinement is on, one `is_final: true, speech_final: false` event is followed by one `is_final: true, speech_final: true`. When refinement is off, every final is `speech_final: true`. See [Canonical-answer semantics](#canonical-answer-semantics-speech-final).
</ResponseField>

<ResponseField name="is_partial" type="boolean">
  `true` for interim results only.
</ResponseField>

<ResponseField name="sentence_id" type="integer">
  Monotonically increasing counter per session.
</ResponseField>

<ResponseField name="duration" type="number">
  Duration (seconds) of the audio segment transcribed.
</ResponseField>

<ResponseField name="latency" type="number">
  Seconds from processing start to result ready (excludes queue time).
</ResponseField>

<ResponseField name="words" type="array">
  Word-level timestamps `[{word, start, end, confidence}]`. Note: the field is **`confidence`**, not `probability`. Present on finals; empty on interims.
</ResponseField>

<ResponseField name="utterance_end_ms" type="integer">
  Timestamp (ms) of last word in the utterance.
</ResponseField>

<ResponseField name="processing_mode" type="string">
  Internal mode string — useful for debugging.
</ResponseField>

### Canonical-answer semantics: `speech_final`

`is_final` and `speech_final` are **NOT identical** when LLM refinement is active — they split into two distinct meanings:

| `is_final` | `speech_final` | Meaning                                                                                                                                                                                                                          |
| ---------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `false`    | `false`        | Interim partial — text may still change as more audio arrives.                                                                                                                                                                   |
| `true`     | `false`        | End-of-speech reached, dict-corrected text. **The LLM is still processing.** A follow-up event with `speech_final: true` will arrive shortly with the canonical text. Only emitted when refinement is active for this utterance. |
| `true`     | `true`         | **Canonical answer.** Definitive, will not be revised. LLM-refined text when context was supplied, otherwise the original ASR text.                                                                                              |

The same `sentence_id` is echoed across both phases so clients can reconcile.

**Canonical event example (after LLM refinement):**

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "transcription",
    "sentence_id": 3,
    "text": "डॉक्टर साहब, मेरा sugar level बहुत high है। Metformin की dose बढ़ाओ।",
    "confidence": 0.87,
    "language": "hi",
    "language_name": "HI",
    "is_final": true,
    "speech_final": true,
    "is_partial": false,
    "duration": 1.82,
    "words": [
      { "word": "डॉक्टर", "start": 0.0, "end": 0.32, "confidence": 0.94 }
    ],
    "speakers": null,
    "timestamp": 1700000000.789
  }
  ```
</ResponseExample>

**Guarantees:**

* **Exactly one canonical event per utterance.** When refinement is on, you get two `transcription` events per utterance (first emit + canonical). When refinement is off, you get one (`speech_final: true`). Never zero, never three.
* **Same `sentence_id`** across both phases. Reconcile on that key.
* **The canonical always arrives.** Consumers waiting on `speech_final: true` never hang.
* **`sentence_id` ordering is preserved per session**, but canonicals are NOT guaranteed to arrive in `sentence_id` order when LLM is on — two utterances finalizing close in time may complete refinement out of order. Key on `sentence_id`, not arrival order.
* **`words[]` corresponds to the original ASR output** on both phases — the LLM does not realign tokens. Use `words[]` for word-level timing, `text` for display.

**Recommended client patterns:**

*Simplest — don't care about the first-emit optimization:*

```js theme={null}
function onMessage(msg) {
  if (msg.type === 'transcription' && msg.speech_final) {
    // Canonical — render and forget
    render(msg);
  }
  // Ignore is_final && !speech_final  (intermediate, will be replaced)
  // Ignore is_partial                 (interim, handle separately if needed)
}
```

*With fast first-emit (UX-aware):*

```js theme={null}
const livePainted = new Map();   // sentence_id → line slot

function onMessage(msg) {
  if (msg.type !== 'transcription') return;
  const sid = msg.sentence_id;

  if (msg.is_partial) { renderPartial(msg); return; }

  let entry = livePainted.get(sid);
  if (!entry) {
    entry = createLine();
    livePainted.set(sid, entry);
  }
  entry.text    = msg.text;
  entry.pending = msg.is_final && !msg.speech_final;   // dim while LLM runs
  render(entry);

  if (msg.speech_final) {
    finalize(entry);
    livePainted.delete(sid);
  }
}
```

**For voicebot NLU routing**: feed the first-emit text (`speech_final: false`) to NLU immediately for fast intent dispatch — don't wait for canonical. If your NLU benefits from proper-noun accuracy (name-spelling slots, drug-name lookup), run a second-pass call on the canonical (`speech_final: true`) text and reconcile on `sentence_id`.

<Note>
  **Legacy `refined` event.** Earlier builds emitted a separate `refined` event \~400 ms after the final instead of a second `transcription`. The 60db `/ws/stt` proxy transparently handles both shapes — if you're still seeing `refined` events in the wire trace, upstream workers haven't been restarted onto the two-phase build yet. New client code should target the two-phase flow only; `refined` is accepted but deprecated.
</Note>

### `language_changed` — After `config` message changes language

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "language_changed",
    "language": "Multi-language: HI",
    "language_code": ["hi"]
  }
  ```
</ResponseExample>

### `mode_changed` — After `config` message changes `continuous_mode`

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "mode_changed",
    "continuous_mode": true,
    "mode_name": "continuous",
    "silence_threshold": 0.5
  }
  ```
</ResponseExample>

### `session_stopped` — After `stop` is processed

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "session_stopped",
    "billing_summary": {
      "total_duration_seconds": 5.2,
      "total_cost": 0.000043,
      "characters_transcribed": 42
    }
  }
  ```
</ResponseExample>

### `error` — Processing error

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "error",
    "error": "Audio processing error: ...",
    "timestamp": 1700000000.0
  }
  ```
</ResponseExample>

### `test_response` — Reply to `test` ping

<ResponseExample>
  ```json Response theme={null}
  {
    "type": "test_response",
    "message": "pong - Sentence-based STT ready",
    "timestamp": 1700000000000,
    "processing_mode": "sentence_based"
  }
  ```
</ResponseExample>

## Complete Example

<Tabs>
  <TabItem value="javascript" label="JavaScript (Node.js)">
    ```javascript theme={null}
    const WebSocket = require('ws');

    const API_KEY = 'sk_live_your_key';
    const ws = new WebSocket(`ws://api.60db.ai/ws/stt?apiKey=${API_KEY}`);

    ws.on('open', () => {
      console.log('✓ Connected');
    });

    ws.on('message', (data) => {
      const msg = JSON.parse(data);
      console.log('←', msg.type || Object.keys(msg)[0]);

      if (msg.connection_established) {
        console.log('  User ID:', msg.connection_established.user_id);
        console.log('  Credits:', msg.connection_established.credit_balance);

        // Start session
        ws.send(JSON.stringify({
          type: 'start',
          languages: ['en', 'hi'],
          config: {
            encoding: 'mulaw',
            sample_rate: 8000,
            continuous_mode: true,
            utterance_end_ms: 500,
            interim_results_frequency: 300
          }
        }));

      } else if (msg.type === 'connected') {
        console.log('✓ Session started! Send audio now...');

        // Send audio chunks (480 bytes = ~60ms at 8kHz)
        const audioInterval = setInterval(() => {
          const audioChunk = getAudioChunk();
          ws.send(audioChunk);
        }, 60);

        // Stop after 5 seconds
        setTimeout(() => {
          clearInterval(audioInterval);
          ws.send(JSON.stringify({ type: 'stop' }));
        }, 5000);

      } else if (msg.type === 'speech_started') {
        console.log('🎤 Speech detected - barge-in opportunity');

      } else if (msg.type === 'transcription') {
        if (msg.is_final) {
          console.log('✓ Final:', msg.text, `(confidence: ${msg.confidence})`);
        } else {
          console.log('  Partial:', msg.text);
        }

      } else if (msg.type === 'session_stopped') {
        console.log('✓ Session stopped');
        console.log('  Duration:', msg.billing_summary.total_duration_seconds, 's');
        console.log('  Cost: $', msg.billing_summary.total_cost);
        ws.close();
      }
    });

    ws.on('error', (error) => {
      console.error('Error:', error);
    });

    ws.on('close', () => {
      console.log('Connection closed');
    });
    ```
  </TabItem>

  <TabItem value="python" label="Python">
    ```python theme={null}
    import asyncio
    import json
    import websockets

    async def stt_websocket():
        API_KEY = "sk_live_your_key"
        url = f"ws://api.60db.ai/ws/stt?apiKey={API_KEY}"

        async with websockets.connect(url) as ws:
            # Wait for connection
            msg = json.loads(await ws.recv())
            if msg.get('connection_established'):
                print(f"✓ Connected (User: {msg['connection_established']['user_id']})")
                print(f"  Credits: ${msg['connection_established']['credit_balance']}")

            # Start session
            await ws.send(json.dumps({
                "type": "start",
                "languages": ["en", "hi"],
                "config": {
                    "encoding": "mulaw",
                    "sample_rate": 8000,
                    "continuous_mode": True,
                    "utterance_end_ms": 500,
                    "interim_results_frequency": 300
                }
            }))

            # Wait for connected
            msg = json.loads(await ws.recv())
            assert msg["type"] == "connected"
            print("✓ Session started!")

            # Send audio for 5 seconds
            for _ in range(83):  # ~5000ms / 60ms
                audio_chunk = get_audio_chunk()
                await ws.send(audio_chunk)  # Send as binary
                await asyncio.sleep(0.06)

            # Stop session
            await ws.send(json.dumps({"type": "stop"}))

            # Process remaining messages
            while True:
                msg = json.loads(await ws.recv())
                msg_type = msg.get("type")

                if msg_type == "speech_started":
                    print("🎤 Speech detected")

                elif msg_type == "transcription":
                    if msg.get("is_final"):
                        print(f"✓ {msg['text']} (confidence: {msg['confidence']})")
                    else:
                        print(f"  {msg['text']}...")

                elif msg_type == "session_stopped":
                    print(f"✓ Session stopped")
                    print(f"  Duration: {msg['billing_summary']['total_duration_seconds']}s")
                    print(f"  Cost: ${msg['billing_summary']['total_cost']}")
                    break

    asyncio.run(stt_websocket())
    ```
  </TabItem>

  <TabItem value="browser" label="Browser">
    ```javascript theme={null}
    const ws = new WebSocket('ws://api.60db.ai/ws/stt?apiKey=sk_live_your_key');
    let mediaRecorder;
    let audioContext;

    ws.onopen = () => {
      console.log('✓ Connected');
    };

    ws.onmessage = (event) => {
      const msg = JSON.parse(event.data);

      if (msg.connection_established) {
        console.log('✓ Authenticated!');

        // Start session
        ws.send(JSON.stringify({
          type: 'start',
          languages: ['en'],
          config: {
            encoding: 'linear',
            sample_rate: 48000,
            continuous_mode: true,
            interim_results_frequency: 300
          }
        }));

      } else if (msg.type === 'connected') {
        console.log('✓ Session started! Start speaking...');
        startAudioCapture();

      } else if (msg.type === 'speech_started') {
        console.log('🎤 Speech detected');

      } else if (msg.type === 'transcription') {
        if (msg.is_final) {
          console.log('✓', msg.text);
          updateTranscriptDisplay(msg.text);
        } else {
          console.log('...', msg.text);
          updateInterimDisplay(msg.text);
        }

      } else if (msg.type === 'session_stopped') {
        console.log('✓ Session stopped');
        stopAudioCapture();
        ws.close();
      }
    };

    async function startAudioCapture() {
      const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
      audioContext = new AudioContext({ sampleRate: 48000 });
      const source = audioContext.createMediaStreamSource(stream);
      const processor = audioContext.createScriptProcessor(4096, 1, 1);

      processor.onaudioprocess = (e) => {
        const inputData = e.inputBuffer.getChannelData(0);
        const pcm16 = new Int16Array(inputData.length);
        for (let i = 0; i < inputData.length; i++) {
          pcm16[i] = Math.max(-32768, Math.min(32767, inputData[i] * 32768));
        }

        ws.send(JSON.stringify({
          type: 'audio',
          audio: btoa(String.fromCharCode(...new Uint8Array(pcm16.buffer))),
          encoding: 'linear',
          sample_rate: 48000
        }));
      };

      source.connect(processor);
      processor.connect(audioContext.destination);
      mediaRecorder = { stream, processor, source };
    }

    function stopAudioCapture() {
      if (mediaRecorder) {
        mediaRecorder.stream.getTracks().forEach(track => track.stop());
        mediaRecorder.processor.disconnect();
        audioContext.close();
      }
    }
    ```
  </TabItem>
</Tabs>

## Audio Requirements

| Property    | Telephony (μ-law) | Browser (PCM)                 |
| ----------- | ----------------- | ----------------------------- |
| Encoding    | `mulaw` (8-bit)   | `linear` (16-bit)             |
| Sample Rate | 8000 Hz           | 16000, 24000, 44100, 48000 Hz |
| Chunk Size  | 480 bytes (60ms)  | 960-1920 bytes (60-120ms)     |
| Channels    | Mono (1 channel)  | Mono (1 channel)              |

## Supported Languages

39 languages total, backed by Parakeet-TDT (25 European), Vaani-FastConformer (13 Indic + Hinglish), and FC-Arabic (MSA). Fetch the full catalog from `GET /stt/languages`.

**Not supported** (explicit rejection, no silent aliasing): `ur`, `ja`, `ko`, `zh`, `th`, `vi`, `id`, `tl`, `sw`, `tr`, `fa`, `he`. Arabic dialect tags (`ar-eg`, `ar-lv`, …) return `dialect_not_supported` — pass `ar` for best-effort MSA.

Common languages:

| Code | Language  | Code | Language |
| ---- | --------- | ---- | -------- |
| `en` | English   | `hi` | Hindi    |
| `bn` | Bengali   | `es` | Spanish  |
| `fr` | French    | `de` | German   |
| `gu` | Gujarati  | `ta` | Tamil    |
| `te` | Telugu    | `kn` | Kannada  |
| `ml` | Malayalam | `mr` | Marathi  |
| `pa` | Punjabi   | `ar` | Arabic   |

## Pricing

* **Rate**: \$0.00000833 per second
* **Minimum**: \$0.01 per session
* **Billing**: Per second of audio processed

## Related

* [TTS WebSocket](/websocket-api/tts) - Text-to-Speech endpoint
* [WebSocket Quick Start](/websocket-api/quickstart) - Get started guide
* [WebSocket Playground](/websocket-playground) - Test in browser
