DataEyesAI
Official SiteConsoleDocs HomeGetting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs HomeGetting StartedDeveloper ToolsAI Models API
  1. Audio API
  • OpenAI format (supports major original models)
    • Chat (Response)
      • Create Network Search
      • Create Model Response GPT-5 Enable Thinking
      • Create Function Call
      • Create Model Response
      • Create Model Response (Streaming Return)
      • Create Model Response (Control Thinking Length)
    • ChatGPT Interface
      • Audio
        • Audio to text gpt-4o-transcribe
        • GPT-4o-audio
        • Audio to text whisper-1
        • Audio to text gpt-4o-transcribe
        • Create voice gpt-4o-mini-tts
      • Chat
        • Create chat-based image recognition (non-streaming)
        • Create chat-based image recognition (streaming)
        • Create chat-based image recognition (streaming) best64
        • Official N test
        • Create structured output
        • Control the effort level of the inference model
        • Create chat function call
        • deepseek-ocr recognition
        • Create chat completion (non-stream)
      • Completions
        • ChatGPT automatic completion
        • Create completion
    • Image
      • Edit image
      • Create chat completion (streaming)
      • Create chat completion (qwen-mt-turbo)
      • Create chat completion with deepseek v3.1 level of reasoning (streaming)
    • Audio
      • Speech recognition
      • Speech synthesis
      • Official Function Calling invocation
      • Create chat-generated images (non-streaming)
    • Embedding
      • Text embeddings
  • Anthropic format
    • Chat
    • Chat(prompt cache)
    • Streaming response
    • Chat (deep reasoning)
    • Tool invocation (function call)
    • Analyze image
  • Google Gemini interface
    • Native format
      • Text-to-image + control over aspect ratio + clarity
      • Generate image
      • Text generation
      • Text generation - stream
      • Text generation + reasoning - stream
      • Image generation
      • Formatted output
      • Function call
      • Document understanding
      • URL context [native format]
      • Code execution
      • Video understanding
      • URL context
      • Video understanding - url [native format]
      • Imagen 4
      • Audio understanding
      • Embeddings
      • Chat
      • Edit image
    • Image-to-image Base64 request method
      • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
      • Image editing
      • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Image generation( gemini-2.5-flash-image)
      • Image generation gemini-2.5-flash-image, controlling aspect ratio.
      • Image understanding
    • Image-to-image URL request returns URL request format OpenAI
      • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Image understanding
  • NanoBanana
    • OpenAI request
      • Edit image
      • OpenAI image format
    • Gemini request
      • Generate image
      • Edit image
  • Midjourney format
    • Midjourney API Reference
    • Task query interface
    • Upload image
    • Get seed (Seed)
    • Submit Imagine task
    • Query tasks based on ID list
    • FaceSwap
    • Execute Action operation
    • /mj/submit/blend
    • Submit Describe task
    • Submit Modal
    • Refresh link
    • Edit image
    • Query task status by task ID
    • Get the seed of the task image
  • Doubao - Painting
    • doubao-seededit-3-0-i2i-250628
    • doubao-seedream-4-0-250828 - text-to-image
    • doubao-seedream-4-0-250828 - image-to-image
    • doubao-seedream-4-0-250828 - multi-image generation
  • Rerank Reordering Model
    • Rerank
  • Video Model
    • Grok Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Video-Editing
      • 05-Video-Extension
    • Seedance Video Generation
      • 00-Overview
      • 01-Create-Video-Generation-Task
      • 02-Query-Video-Generation-Task
      • 03-Query-Video-Generation-Task-List
      • 04-Cancel-or-Delete-Task
      • Seedance Private Asset Library API Documentation
    • MiniMax-H3 Video Generation
      • 00-Overview
      • 01-Create-Video-Generation
      • 02-Create-Video-Regeneration
      • 03-Create-H3-Context-IR
      • 04-Query-Task
      • 05-List-Tasks
      • 06-Cancel-or-Delete-Task
    • Hailuo Video Generation
      • 00-Overview
      • 01-Text-to-Video-T2V
      • 02-Image-to-Video-I2V
      • 03-First-Last-Frame-FL2V
      • 04-Subject-Reference-S2V
      • 05-Query-Task-Status
      • 06-Video-Download
      • 99-Appendix-Camera-Movement-and-Webhooks
    • Jimeng Video Generation
      • 00-Overview
      • 01-3.0-Pro-Video-Generation
      • 02-720P-Text-to-Video
      • 03-720P-Image-to-Video-First-Frame
      • 04-720P-Image-to-Video-Start-End-Frame
      • 05-720P-Image-to-Video-Camera
      • 06-1080P-Text-to-Video
      • 07-1080P-Image-to-Video-First-Frame
      • 08-1080P-Image-to-Video-Start-End-Frame
      • 09-Error-Codes
    • Kling AI Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Omni-Video
      • 04-Multi-Image-to-Video
      • 05-Motion-Control
      • 06-Multi-Elements
      • 07-Video-Extension
      • 08-Lip-Sync
      • 09-Avatar
      • 10-Text-to-Audio
      • 11-Video-to-Audio
      • 12-TTS
      • 13-Custom-Voices
      • 14-Image-Recognition
      • 15-Element-Management
      • 16-Video-Effects
    • Vidu Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Start-End-Frame
      • 05-Multi-Frame
      • 06-Scene-Template
      • 07-Template-Story
      • 08-Query-Tasks
    • HappyHorse
      • HappyHorse Text-to-Video
      • HappyHorse Image-to-Video (First Frame)
      • HappyHorse Reference-to-Video
      • HappyHorse Video Editing
    • Wan Video Generation
      • 00-Overview.md
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Video-Editing
      • 05-First-Last-Frame-to-Video
      • 06-Motion-Transfer-and-Character-Swap
      • 07-Digital-Human-Video
      • 08-VACE-Video-Editing
      • 09-Query-Task
  • Audio API
    • Audio API
    • Gemini TTS API
    • Google DeepMind Lyria API
    • Elevenlabs Speech to Text API Reference
    • Text-to-Music Suno
      • Task Submission
        • Generate Song (Inspiration Mode)
        • Generate Song (Custom Mode)
        • Generate Song (Continuation Mode)
        • Generate Song (Singer Style)
        • Generate Song (Secondary Creation from Uploaded Song)
        • Generate Song (Song Stitching)
        • Generate Lyrics
        • Song Stitching
      • Query Interface
        • Batch Retrieve Tasks
        • Query Single Task
  1. Audio API

Elevenlabs Speech to Text API Reference

Version: v1.0.0  |  Last Updated: 2026-07-15  |  Models: scribe_v2, scribe_v2_realtime  |  Status: GA
This document is written from live requests/responses against the production environment (https://platform.dataeyes.ai). All field names, types, and values strictly match actual API behavior.

Table of Contents#

1. Overview
2. Authentication
3. Available Models & Pricing
4. File Transcription (HTTP)
4.1 Endpoint
4.2 Request Parameters
4.3 Response Body
4.4 The words Array
5. Realtime Transcription (WebSocket)
5.1 Connection URL
5.2 Query Parameters
5.3 Client Messages
5.4 Server Messages
5.5 Audio Format Requirements
6. Complete Examples
6.1 cURL — File Transcription
6.2 Python — File Transcription
6.3 Python — Realtime (WebSocket)
6.4 Node.js — Realtime (WebSocket)
7. Billing
8. Error Codes
9. Best Practices
10. Changelog

1. Overview#

The Speech to Text API provides high-accuracy speech recognition powered by the Scribe v2 model family, covering two typical scenarios:
ScenarioModelProtocolBest for
File transcriptionscribe_v2HTTP (multipart/form-data)Offline/batch transcription of complete audio files: meeting recordings, podcasts, call QA, etc.
Realtime streamingscribe_v2_realtimeWebSocketLow-latency transcribe-as-you-speak: live captions, voice input, meeting notes, etc.
Key capabilities:
CapabilityDescription
Multilingual90+ languages with automatic language detection and confidence score
Word-level timestampsMillisecond-precision start / end per word; the realtime model additionally returns character-level timestamps
Speaker diarizationMulti-speaker audio automatically labeled with speaker_id
Confidence scoresEach word carries logprob (log probability) usable for quality filtering
Incremental streaming outputWebSocket sessions continuously push partial_transcript (tentative) and committed_transcript (final) results

2. Authentication#

All requests must carry an API Key:
Authorization: Bearer YOUR_API_KEY
ProtocolHow to send
HTTPRequest header Authorization: Bearer <API_KEY>
WebSocketHandshake request header Authorization: Bearer <API_KEY>
Security note: Your API Key is a sensitive credential. Never expose it in client-side code, public repositories, or logs. Inject it via environment variables or a secrets manager.

3. Available Models & Pricing#

Model IDDescriptionProtocolBillingStatus
scribe_v2File transcription model. High accuracy with speaker diarization and word-level timestamps.HTTPBy audio durationGA
scribe_v2_realtimeRealtime streaming transcription model. Low-latency incremental output with word- and character-level timestamps.WebSocketBy audio durationGA
Both models are billed by audio duration (pay-as-you-go), independent of request count — see 7. Billing for metering rules. Refer to the console model-pricing page for the authoritative live unit price.

4. File Transcription (HTTP)#

4.1 Endpoint#

POST https://platform.dataeyes.ai/v1/elevenlabs/speech-to-text
Request format: multipart/form-data.

4.2 Request Parameters#

ParameterTypeRequiredDescription
model_idstringYesModel identifier. Currently available: scribe_v2. Missing value returns 400.
filefileOne of twoAudio file (binary upload). At least one of file / source_url is required. Common formats supported: wav / mp3 / m4a / flac / ogg / webm, etc.
source_urlstringOne of twoPublic URL of the audio file. Must be directly reachable by the transcription servers (private/intranet URLs or some region-restricted CDNs may fail). Prefer file upload.
language_codestringNoAudio language (ISO 639-1/639-3, e.g. en, zh). Auto-detected if omitted. Setting it explicitly skips detection and improves accuracy (measured: language_probability becomes 1.0).
diarizebooleanNoEnable speaker diarization, default false. When enabled, each item in words[] carries a speaker_id (e.g. speaker_0, speaker_1).
num_speakersintegerNoHint for the maximum number of speakers; use together with diarize=true to improve separation accuracy.
timestamps_granularitystringNoTimestamp granularity: word (default) / character / none.
tag_audio_eventsbooleanNoTag non-speech audio events (laughter, applause, etc.), default true.
All other official ElevenLabs Speech-to-Text parameters (e.g. additional_formats, file_format) are passed through verbatim and behave as officially documented.

4.3 Response Body#

Returns HTTP 200 on success. Top-level fields:
FieldTypeDescription
language_codestringDetected/specified language code (ISO 639-3), e.g. "eng".
language_probabilityfloatLanguage-detection confidence, 0–1. Equals 1.0 when language_code was explicitly provided.
textstringFull transcript text.
wordsarray<object>Per-word details with timestamps and confidence — see 4.4 The words Array.
transcription_idstringUnique identifier of this transcription, useful for troubleshooting and log tracing.
audio_duration_secsfloatTotal duration of the audio file in seconds. The billed duration is based on the end time of the last transcribed word — see 7. Billing.

Sample Response#

Sample response (9.06-second English test audio, words truncated):
{
  "language_code": "eng",
  "language_probability": 0.8746970891952515,
  "text": "Hello DataEyes, this is a speech-to-text API test. The quick brown fox jumps over the lazy dog",
  "words": [
    {
      "text": "Hello",
      "start": 0.219,
      "end": 0.5,
      "type": "word",
      "logprob": -0.0000160931
    },
    {
      "text": " ",
      "start": 0.5,
      "end": 0.599,
      "type": "spacing",
      "logprob": -0.0573769733
    },
    {
      "text": "DataEyes,",
      "start": 0.599,
      "end": 1.179,
      "type": "word",
      "logprob": -0.2447926141
    }
  ],
  "transcription_id": "wnnnNJ3sHGJrl8gXgXyS",
  "audio_duration_secs": 9.06
}

4.4 The words Array#

FieldTypeDescription
textstringWord text (punctuation included). A space when type=spacing.
startfloatStart time in seconds.
endfloatEnd time in seconds.
typestringElement type: word / spacing (inter-word gap) / audio_event (requires tag_audio_events=true).
logprobfloatLog probability of the word (≤ 0; closer to 0 means higher confidence).
speaker_idstringSpeaker label (e.g. "speaker_0"). Only present when diarize=true.
Parsing tip: Use the top-level text field for plain text; when processing words, filter by type — do not assume word and spacing strictly alternate.

5. Realtime Transcription (WebSocket)#

5.1 Connection URL#

wss://platform.dataeyes.ai/v1/elevenlabs/realtime
A successful handshake returns 101 Switching Protocols. Authentication failures are rejected at the handshake with HTTP 401 — no connection is established.

5.2 Query Parameters#

ParameterTypeRequiredDescription
model_idstringNoModel identifier, default scribe_v2_realtime (currently the only available value; passing it explicitly is recommended).
audio_formatstringNoAudio encoding, default pcm_16000. Format is pcm_<sample_rate> and must match the audio you actually send — see 5.5.
language_codestringNoAudio language. Auto-detected if omitted.
commit_strategystringNoCommit strategy. Set to vad to let server-side voice-activity detection segment and commit automatically; by default the client commits manually via the commit field.
Others—NoAll other official ElevenLabs Realtime query parameters are passed through verbatim.
Note: include_timestamps is force-set to true by the platform so that sessions can be billed precisely from word-level timestamps at session end.

5.3 Client Messages#

After the connection is established, the client sends audio chunks as JSON text frames:
{
  "message_type": "input_audio_chunk",
  "audio_base_64": "<Base64-encoded PCM audio data>",
  "sample_rate": 16000,
  "commit": true
}
FieldTypeRequiredDescription
message_typestringYesFixed value "input_audio_chunk".
audio_base_64stringYesBase64-encoded raw PCM audio (16-bit, mono, little-endian, no file header).
sample_rateintegerNoSample rate; must match audio_format, e.g. 16000.
commitbooleanNoWhen true, the server commits the audio accumulated so far and produces final transcript results. When streaming continuously, send chunks and set true on the last one.

5.4 Server Messages#

The server pushes JSON text frames in the following order:
session_started → partial_transcript (0–N times) → committed_transcript → committed_transcript_with_timestamps

session_started — session established#

{
  "message_type": "session_started",
  "session_id": "b4c1d85547624f98adecc5a4df82511a",
  "config": {
    "sample_rate": 16000,
    "audio_format": "pcm_16000",
    "language_code": null,
    "timestamps_granularity": "word",
    "vad_commit_strategy": false,
    "vad_silence_threshold_secs": 1.5,
    "vad_threshold": 0.4,
    "min_speech_duration_ms": 100,
    "min_silence_duration_ms": 100,
    "max_tokens_to_recompute": 5,
    "model_id": "scribe_v2_realtime",
    "include_timestamps": true,
    "include_language_detection": false,
    "filter_background_audio": false,
    "keyterms": [],
    "no_verbatim": false,
    "entity_detection": null
  }
}
FieldDescription
session_idUnique session identifier for troubleshooting and reconciliation.
configThe effective session configuration (query parameters merged with defaults). Validate it before sending audio.

partial_transcript — tentative transcript (incremental, optional)#

Low-latency intermediate result; may be revised by subsequent audio. Suitable for live-caption previews:
{ "message_type": "partial_transcript", "text": "Hello, DataEyes. This is" }

committed_transcript — final transcript#

Returned after a commit; the text will no longer change:
{
  "message_type": "committed_transcript",
  "text": "Hello, DataEyes. This is a speech-to-text API test. The quick brown fox jumps over the lazy dog."
}

committed_transcript_with_timestamps — final transcript with timestamps#

Immediately follows committed_transcript, carrying full word- and character-level timestamps:
{
  "message_type": "committed_transcript_with_timestamps",
  "text": "Hello, DataEyes. This is a speech-to-text API test. The quick brown fox jumps over the lazy dog.",
  "language_code": null,
  "words": [
    {
      "text": "Hello,",
      "start": 0.219,
      "end": 0.479,
      "type": "word",
      "speaker_id": null,
      "logprob": -0.5185597737,
      "characters": [
        { "text": "H", "start": 0.219, "end": 0.239 },
        { "text": "e", "start": 0.239, "end": 0.319 },
        { "text": "l", "start": 0.319, "end": 0.34 },
        { "text": "l", "start": 0.34, "end": 0.36 },
        { "text": "o", "start": 0.36, "end": 0.479 },
        { "text": ",", "start": 0.479, "end": 0.479 }
      ],
      "channel_index": null
    }
  ]
}
The words[] structure matches the HTTP endpoint (see 4.4) with two additional fields:
FieldTypeDescription
charactersarray<object>Character-level timestamps; each item has text / start / end.
channel_indexinteger | nullAudio channel index; null for mono audio.

Error events#

Upstream service errors are relayed to the client as messages (such sessions incur no charge):
message_typeMeaning
scribeAuthErrorService authentication error
scribeQuotaExceededErrorService quota exceeded
scribeThrottledErrorService throttled
scribeSessionTimeLimitExceededErrorSession exceeded the maximum duration limit

5.5 Audio Format Requirements#

ItemRequirement
EncodingRaw PCM (16-bit signed, little-endian), no WAV/RIFF header
ChannelsMono
Sample rateMust match the audio_format parameter; pcm_16000 (16 kHz) recommended
Extract PCM from a WAV file (ffmpeg):

6. Complete Examples#

6.1 cURL — File Transcription#

6.2 Python — File Transcription#

6.3 Python — Realtime (WebSocket)#

Dependency: pip install websockets

6.4 Node.js — Realtime (WebSocket)#

Dependency: npm install ws

7. Billing#

Both models are billed by actual audio duration, independent of request count or output text length. Refer to the console model-pricing page for the current unit price.
Metering rules:
The billed duration is the audio span covered by the transcript — the end time of the last word, rounded up to whole seconds (trailing silence is not charged).
Duration is metered in audio tokens: 1 minute of audio = 1,000 audio tokens (partial minutes prorated and rounded up), reported as prompt_tokens in the usage logs.
Example: 9.06 s audio with last word end = 8.319s → rounded to 9 s → 9 / 60 × 1000 = 150 audio tokens.
Realtime sessions are settled when the connection closes; sessions that produce no committed transcript (e.g. connect then disconnect immediately) cost 0.
Failed request validation (400) and failed authentication (401) are never charged; WebSocket sessions that receive an upstream error event are not charged.
Per-request audio-token usage and cost details are visible in the console logs.

8. Error Codes#

HTTP statusScenarioError message exampleHandling
400Missing required parameter such as model_idmodel_id is requiredComplete the required fields per 4.2.
401API Key missing, malformed, or revoked (applies to both HTTP and WebSocket handshake)无效的令牌 (invalid token)Confirm the header is Authorization: Bearer <key> and the key is valid.
403The key has no access to the requested model—Confirm the API Key has access to the requested model.
429Rate limit exceeded or insufficient quota—Reduce request rate and retry with exponential backoff; upgrade the plan if needed.
500Invalid request (e.g. neither file nor source_url provided) or internal server erroreither file or source_url is requiredCheck request completeness first; for transient server errors, contact support with the request id from the error message.
503Service temporarily unavailable or overloaded—Retry with exponential backoff.
All error responses share one JSON structure: {"error": {"code": "...", "message": "... (request id: ...)", "type": "..."}}. Include the request id from message when reporting issues.

9. Best Practices#

9.1 Choosing a model#

You already have a complete audio file → use scribe_v2 (HTTP): one request, accuracy first.
You need transcribe-as-you-speak → use scribe_v2_realtime (WebSocket): render partial_transcript for low-latency previews and treat committed_transcript_with_timestamps as the final result.

9.2 HTTP transcription#

Prefer binary file upload. source_url requires the transcription servers to reach the URL directly — intranet or restricted-CDN URLs will fail to fetch.
Pass language_code explicitly when the language is known: it skips detection and improves accuracy.
Transcription time grows with audio length; scale client timeouts with duration (start at ≥ 300 s).

9.3 Realtime sessions#

Audio must be raw header-less PCM (16-bit / mono) with a sample rate matching audio_format — mismatched formats produce empty or garbled transcripts.
In production, stream audio in 100–500 ms chunks and set commit: true at the end of an utterance (or use commit_strategy=vad for server-side segmentation).
Validate the config returned in session_started to confirm your parameters took effect.
Close the connection only after receiving committed_transcript_with_timestamps to get complete timestamps; settlement happens on close.
Record session_id (WebSocket) and transcription_id (HTTP) for troubleshooting and usage reconciliation.

9.4 Working with results#

Use the top-level text for plain transcripts; when processing words, branch on the type field (word / spacing / audio_event).
logprob closer to 0 means higher confidence; flag low-confidence words (e.g. logprob < -1.0) for human review.
For diarized audio, group words by speaker_id to reconstruct the per-speaker dialogue.

10. Changelog#

DateVersionChanges
2026-07-15v1.0.0Initial release. Supports scribe_v2 (HTTP file transcription) and scribe_v2_realtime (WebSocket realtime transcription).

Written from live production testing (platform.dataeyes.ai) on 2026-07-15. Test sample: 9.06-second 16 kHz mono English audio; HTTP transcription returned 33 words elements; the realtime session message flow session_started → committed_transcript → committed_transcript_with_timestamps was verified end to end.
Previous
Google DeepMind Lyria API
Next
Generate Song (Inspiration Mode)