Skip to content

1. Model Description

ability Seeduplex Full-Duplex Real-Time Voice
End-to-end real-time voice dialogue (low-latency, full-duplex) ✅
Streaming ASR / Streaming Text Response / Streaming Audio Synthesis ✅
Function Calling ✅
Context Injection and Management ✅
Proprietary Capabilities (extension: ASR/TTS/Dialog passthrough, e. g., internet access, singing, hot words, etc.) Functions such as internet connection and trending keywords are not enabled for the time being.

2. Details of Functional Interfaces

2.1 Full-Duplex Real-Time Voice WebSocket

2.1.1 Request URL

wss://genaiapi-m2.cloudsway.net/ws/api/v1/ai/{YOUR_ENDPOINT}/doubao/s2s
  • {YOUR_DOMAIN}: The access domain name assigned by the platform.

  • {YOUR_ENDPOINT}: The access endpoint identifier (path segment) assigned by the platform.

2.1.2 Handshake Request Header Parameters

Parameter Name Field Type Required or not default value Description
Authorization string Either one of the two -
API Key Credential (the authentication header agreed with the platform)

2.1.3 Message Frames and Common Fields

  • Each WebSocket Text Message carries one complete event JSON.

  • The sender and receiver shall perform distribution processing according to the top-level typefield.

  • The following fields are reused across multiple events:

Parameter Name Field Type Required or not Description
type string Yes Event type, see §2.2 and §2.3
event_id string No When a Client initiates a request, it can generate and pass a custom event_id, which is an optional field but recommended to be provided for subsequent event matching and tracing.
Client                        Platform
  │── WebSocket connect (auth) ──▶│
  │── session.create ──────────▶│
  │◀── session.created ─────────│   (contains session.id, can continue history)
  │── input_audio_buffer.append ▶│   (loop: audio / other uplink events)
  │◀── transcription / text / audio / usage ──│
  │── session.close ───────────▶│
  │◀── session.closed ──────────│
  │── close WebSocket ──────────▶│

Key Points:

  1. It is recommended that after calling session. create, wait for the session. createdevent before starting stream pushing.

  2. Context continuation: In the session. update, postback the session. idreturned by the previous session. createdin the session. idfield (i. e. dialog id); the server retains the latest 20 rounds of Q\&A by default.

  3. Graceful shutdown: First send session. closeand receive session. closedbefore closing the WebSocket; a direct disconnection may trigger ContextCanceled (error code 55000001).

  4. Keep-alive: input_audio_mute. commit only ensures that the session is maintained with the model side, and does not guarantee that the connection with the platform is sustained. It is recommended to send a mute packet every 10 seconds to maintain the connection with the platform; if no mute packet is received for more than 90 seconds, the platform will release the connection. If there is no interaction for more than 10 minutes, the model side may release the connection (error code 45000003).

  5. In scenarios such as instance maintenance, the Client may first receive the session. closing (see §2.3.21) delivered by the platform, and then the connection will be closed. It is recommended to reconnect after receiving this event to avoid server-side interruption.

2.1.5 Request Example (send session. create after connection establishment)

{
  "type": "session.create",
  "event_id": "evt_create_001",
  "session": {
    "instructions": "You are a friendly voice assistant.",
    "audio": {
      "input": { "format": { "type": "pcm", "rate": 16000 } },
      "output": {
        "format": { "type": "ogg_opus", "rate": 24000 },
        "voice": "YOUR_VOICE_ID",
        "speed": 0,
        "loudness": 0
      }
    },
    "extension": {
      "asr": {},
      "tts": {},
      "dialog": {}
    }
  }
}

2.1.6 Response Example (session. created)

The response header will contain the X-Tt-Logid field

{
  "type": "session.created",
  "session": {
    "id": "dlg_xxxxxxxx"
  }
}

In the following sections, "request parameters" refer to the fields of the root JSON object of a single frame (excluding the common fields already listed).

2.2.1 session.create / session.update

Parameter Name Field Type Required or not default value Description
type string Yes - Specifies the type of the requested event. When creating a session, this field is fixed as session. create; when updating a session, this field is fixed as session. update
session object Yes - Session Configuration
session.id string No - Corresponds to the original dialog ID, which is passed in when continuing a historical conversation
session.instructions string No - System prompt (System message), which is used to guide the model's response content and audio style, and the total length of it together with the model's internal SP and context has an upper limit of 12K tokens
session.audio object Yes - Audio Input/Output Specifications
session.audio.input object Yes - Input Audio Configuration
session.audio.input.format object No - Specify the specifications for the uploaded audio
session.audio.input.format.type string No - Specify the input audio format, supported formats include pcmand speech_opus
session.audio.input.format.rate int No - Specify the input audio sample rate; only 16000is supported, unit: Hz
session.audio.output object Yes - Output Audio Configuration
session.audio.output.format object No - Specify the output audio specifications
session.audio.output.format.type string No - Specify the output audio format, supported formats include pcmand ogg_opus
session.audio.output.format.rate int No - Specify the output audio sampling rate, which only supports 24000, in Hz
session.audio.output.voice string Yes - Specify the voice ID. For the currently supported voices, please refer to: Voice List; cloned voices are also supported.
session.audio.output.speed
number No 0 Specify the speech rate; the default value is 0, with a value range of [-50, 100]; the larger the value, the faster the speech rate. -50corresponds to 0.5x speed, 100corresponds to 2.0x speed, and the speech rate is not adjusted by default.
session.audio.output.loudness number No 0 Specify the volume, with a default value of 0, and the value range is [-50, 100]; the larger the value, the higher the volume. -50corresponds to 0.5x volume, 100corresponds to 2.0x volume, and the volume will not be adjusted by default.
session.tools array No - Function Calling tool definition, using standard JSON Schema structure
session.extension object No - Configure model passthrough parameters
session.extension.asr object No - Configure ASR parameters
session.extension.asr.extra object No - Configure additional ASR parameters
session.extension.asr.extra.enable_asr_twopass bool No false Enable non-streaming model recognition capability, defaulting to false
session.extension.asr.extra.boosting_table_id string No - Hot word list ID is not supported yet
session.extension.asr.extra.boosting_table_name string No - Name of popular word list not supported yet
session.extension.asr.extra.regex_correct_table_id string No - Not yet supported
session.extension.asr.extra.regex_correct_table_name string No - Not yet supported
session.extension.asr.extra.context string No - Not yet supported
session.extension.dialog object No - Configure Dialog parameters
session.extension.dialog.location string No - Location information configuration supports the following fields: longitude (longitude), latitude (latitude), city (city), country (country), province (province), district (district/county), town (town/township/sub-district), country_code (country code), address (detailed address)
session.extension.dialog.dialog_context array No - Initialize the context by passing in QA pairs in the order of user/ assistantin pairs
session.extension.dialog.dialog_context[].role string No - Specify the dialogue role: user/ assistant
session.extension.dialog.dialog_context[].text string No - Input dialogue text
session.extension.dialog.dialog_context[].timestamp int No - Pass in the dialogue timestamp
session.extension.dialog.extra string No - Configure additional parameters for Dialog
session.extension.dialog.extra.strict_audit bool No true Enable strict review, default value is true
session.extension.dialog.extra.audit_response string No - Customized reply script when a specified user's query triggers security review
session.extension.dialog.extra.enable_volc_websearch bool No false The built-in networking function is not supported yet, and it is set to falseby default. To activate the service, please refer to Console Integrated Information Search API
session.extension.dialog.extra.volc_websearch_type string No - Not yet supported. Used to specify the type of search service, supported values include web_custom_api, web_global_api
session.extension.dialog.extra.enable_music bool No false Enable the singing function, which defaults to false. After being enabled, the system will retrieve singing data from the music library and feed it into the model to improve the model's singing performance
session.extension.dialog.extra.enable_loudness_norm bool No false Enable the audio loudness equalization function for output, which is by default false
session.extension.dialog.extra.enable_user_query_exit bool No false Enable the exit intent detection capability, which is by default false. After being enabled, the server will carry an exit intent signal in the response.output_audio.doneevent for the Client to perform the exit operation
session.extension.tts object No - Configure TTS parameters
session.extension.tts.extra object No - Configure additional TTS parameters
session.extension.tts.extra.max_length_to_filter_parenthesis int No 0 Specify the length of the text within the parentheses to be filtered, in characters. The default value is 0 (meaning no filtering), and the recommended value range is 0 to 100.
session.extension.tts.extra.explicit_dialect string No - Specify the dialect parameter. Supported dialects: dongbei, sichuan, shaanxi, yue, beijing, henan, tianjin, shanghai
session.extension.tts.extra.aigc_metadata string No - Configure AIGC content traceability and copyright metadata

session. updateperforms full coverage on toolsrather than incremental update, supporting dynamic addition and removal of tools.

2.2.2 session.close

Parameter Name Field Type Required or not default value Description
type string Yes - Indicates the type of the request event. When ending a session, this field is fixed as session. close

2.2.3 input_audio_buffer.append

Parameter Name Field Type Required or not default value Description
type string Yes - Specifies the type of request event. When sending a request, this field is fixed as input_audio_buffer. append
audio string Yes - Pass in the audio data, fill the audiofield with the Base64-encoded audio data, and send the complete event in the form of a JSON text frame

2.2.4 input_audio_buffer.commit

Parameter Name Field Type Required or not default value Description
type string Yes - Specifies the type of request event. When sending a request, this field is fixed as input_audio_buffer. commit . After confirming that the audio query has been fully sent, you can send this event to force the model to stop generation.

2.2.5 input_audio_mute.commit / input_audio_unmute.commit

Parameter Name Field Type Required or not default value Description
type string Yes - Sent after the microphone is muted: input_audio_mute. commit; sent when the microphone is unmuted: input_audio_unmute. commit (i. e., set typeto the corresponding event name)

2.2.6 speech_text_buffer.commit

Parameter Name Field Type Required or not default value Description
type string Yes - Specifies the type of the request event. Used for greeting, the field is fixed as speech_text_buffer. commit
text string Yes - Enter the "greeting" text to be synthesized

2.2.7 speech_text_buffer.replacement.append / speech_text_buffer.replacement.commit

Parameter Name Field Type Required or not default value Description
type string Yes - Specifies the type of request event, which is applicable to scenarios such as streaming greeting, or cases where no model chat result is required and text-to-audio synthesis is expected to be specified directly. speech_text_buffer.replacement.append: upload the text to be synthesized in streaming mode;speech_text_buffer.replacement.commit: end packet (text upload completed)
text string Yes - Enter the text to be synthesized

2.2.8 conversation.item.create (Context / Function Calling Result)

Context Initialization

Parameter Name Field Type Required or not default value Description
type string Yes - Specifies the type of request event. When sending a request, this field is fixed as conversation.item.create
items array Yes - Add contextual dialogue information, which can be used to initialize the historical context. A maximum of 20 rounds (40 entries) of complete Q\&A can be submitted at a time.
items[].id string No - Specify a custom contextual conversation ID
items[].type string Yes - Specify the context type, which is fixed as message
items[].role string Yes - Specify the conversation role, supported values user, assistant
items[].content array Yes - Specify the dialogue content

Function Calling Result postback (after response.function_call_arguments.done)

Parameter Name Field Type Required or not default value Description
type string Yes - Type of the request event. When sending a request, this field is fixed as conversation.item.create. This event is used to report the tool execution result after the downstream event response.function_call_arguments.doneis returned, so that the model can continue to generate a response based on the tool data.
items array Yes - Specify function postback information
items[].call_id string Yes - The unique identifier for a function call, used to postback the function call result, must match the call_id delivered by the downstream event response.function_call_arguments.done call_idconsistently
items[].role string Yes - Specify the entry role information, with a fixed value of tool, which identifies this item as a function call result
items[].content array Yes - the result information passed into the function call
items[].content[].type string Yes - Content type, fixed value: input_text
items[].content[].text string Yes - the execution result passed into the function call

2.2.9 conversation.item.update

Parameter Name Field Type Required or not default value Description
type string Yes - Indicates the type of request event. When sending a request, this field is fixed as conversation.item.update
items array Yes - Specify the context information to be updated
items[].id string Yes - Specify the ID of the conversation to be updated
items[].content array Yes - Input the dialogue content to be updated

2.2.10 conversation.item.retrieve

Parameter Name Field Type Required or not default value Description
type string Yes - Request event type. When sending a request, this field is fixed as conversation.item.retrieve
items array No - Specify the context information to be queried
items[].id string No - Pass in the conversation ID to be queried, and the context information of this round will be returned; if no ID is passed in, the complete context information of the latest 20 rounds will be returned

2.2.11 conversation.item.delete

Parameter Name Field Type Required or not default value Description
type string Yes - Request event type. When sending a request, this field is fixed as conversation.item.delete
items array Yes - Specify the context information to be deleted
items[].id string Yes - Specify the ID of the conversation to be deleted

Context deletion is performed in units of dialogue turns. When a userside ID is passed in, the paired assistantreplies will be deleted simultaneously, and vice versa.

2.2.12 response.cancel

Parameter Name Field Type Required or not default value Description
type string Yes - Request event type. When sending a request, this field is fixed as response. cancel. It is used to cancel an ongoing response, meaning the Client actively interrupts the server's broadcast to facilitate the next recognition.

After a successful handshake, the response is also a WebSocket Text Message JSON. Except for the following fields, most events may carry event_id (corresponding to the uplink one or generated by the server).

2.3.1 session.created

Parameter Name Field Type Description
type string Fixed to session. created
session object Session Information
session.id string This event indicates that the session has been successfully started, and the returned session. id (corresponding to the original version's dialog. id) can be used to continue the historical conversation content.

2.3.2 session.updated

Parameter Name Field Type Description
type string Fixed as session. updated
session object This event is session. updatethe acknowledgment (ack) response corresponding to the request, indicating that the session configuration has been successfully updated

2.3.3 session.closed

Parameter Name Field Type Description
type string Fixed to session. closed
— — This event is a response indicating that the session has ended (session. closeresponse)

2.3.4 input_audio_buffer.committed

Parameter Name Field Type Description
type string Fixed to input_audio_buffer. committed
— — This event is used to notify the Client that the input audio buffer has been successfully submitted, marking the end of a user audio input session

2.3.5 conversation.item.input_audio_transcription.started

Parameter Name Field Type Description
type string Fixed as conversation.item.input_audio_transcription.started
— — The model returns when it identifies the first word in the audio stream

2.3.6 conversation.item.input_audio_transcription.delta

Parameter Name Field Type Description
type string Fixed as conversation.item.input_audio_transcription.delta
delta string User's speech text content identified in real time by the model (streaming increment)

2.3.7 conversation.item.input_audio_transcription.completed

Parameter Name Field Type Description
type string Fixed as conversation.item.input_audio_transcription.completed
— — Return when the model determines that the user has finished speaking

2.3.8 conversation.item.input_audio_transcription.failed

Parameter Name Field Type Description
type string Fixed as conversation.item.input_audio_transcription.failed
— — ASR recognition failed

2.3.9 response.output_text.delta

Parameter Name Field Type Description
type string Fixed as response.output_text.delta
delta string Text content of the model's response (streaming increment)

2.3.10 response.output_text.done

Parameter Name Field Type Description
type string Fixed as response.output_text.done
— — Model response text generation completed

2.3.11 response.output_audio.started

Parameter Name Field Type Description
type string Fixed to response.output_audio.started
— — An audio synthesis round starts

2.3.12 response.output_audio.delta

Parameter Name Field Type Description
type string Fixed to response.output_audio.delta
delta string The returned streaming audio data block, encoded in Base64

2.3.13 response.output_audio.done

Parameter Name Field Type Description
type string Fixed as response.output_audio.done
status_code string The first round of audio synthesis by the model is completed. Among them, status_code="20000002"indicates that the model has recognized the user's exit intention.

2.3.14 conversation.item.added

Parameter Name Field Type Description
type string Fixed as conversation.item.added
items array Add acknowledgement (ack) for context requests, and return the context array created successfully

2.3.15 conversation.item.retrieved

Parameter Name Field Type Description
type string Fixed to conversation.item.retrieved
items array Acknowledge (ack) of the query context request, returning the queried context content

2.3.16 conversation.item.deleted

Parameter Name Field Type Description
type string Fixed as conversation.item.deleted
items array Acknowledge (ack) the context deletion request and return the deleted context content
status_code string If there is no context to be deleted, a result such as 40000010 may be returned.
message string Error description, e. g. empty conversation deleted messages

2.3.17 response.function_call_arguments.done

Parameter Name Field Type Description
type string Fixed as response.function_call_arguments.done
items array FC function call parameter generation completed. Downlink itemseach function call item carries a unique call_id, function name nameand the generated parameters arguments (JSON string). After the Client executes the local function, it is required to conversation.item.create (role= tool) postback the result, and carry the same call_id in the postback item

2.3.18 response.done

Parameter Name Field Type Description
type string Fixed to response. done
response object Current round response envelope
response.usage object One round of interaction is completed, and the usage statistics for this time are returned.
response.usage.total_tokens int Total tokens in this round
response.usage.input_tokens int Input token count
response.usage.output_tokens int Number of output tokens
response.usage.input_token_details object Input Token Details
response.usage.input_token_details.text_tokens int Input text token
response.usage.input_token_details.audio_tokens int Input Audio Token
response.usage.input_token_details.image_tokens int Input image tokens
response.usage.input_token_details.cached_tokens int Total number of cache hit tokens
response.usage.input_token_details.cached_tokens_details object Cache Hit Details
response.usage.input_token_details.cached_tokens_details.text_tokens int Cache hit text token
response.usage.input_token_details.cached_tokens_details.audio_tokens int cache hit audio token
response.usage.input_token_details.cached_tokens_details.image_tokens int Cache hit image token
response.usage.output_token_details object Output token details
response.usage.output_token_details.text_tokens int output text token
response.usage.output_token_details.audio_tokens int output audio tokens

Response Example

{
  "type": "response.done",
  "event_id": "event_119",
  "response": {
    "usage": {
      "total_tokens": 796,
      "input_tokens": 246,
      "output_tokens": 550,
      "input_token_details": {
        "text_tokens": 0,
        "audio_tokens": 246,
        "image_tokens": 0,
        "cached_tokens": 0,
        "cached_tokens_details": {
          "text_tokens": 0,
          "audio_tokens": 0,
          "image_tokens": 0
        }
      },
      "output_token_details": {
        "text_tokens": 185,
        "audio_tokens": 365
      }
    }
  }
}

2.3.19 response.canceled

Parameter Name Field Type Description
type string Fixed to response. canceled
— — Client interrupt request response. cancelacknowledgement (ack)

2.3.20 error

Parameter Name Field Type Description
type string Fixed as error
status_code string Error code (may also be located in a nested errorobject, subject to the actual message)
message string Error Description

Error events. For the error list, please refer to the error code instructions in the Access Guide. The Client shall uniformly capture and log all downstream errorevents.

2.3.21 session. closing (Platform Prompt)

Parameter Name Field Type Description
type string Fixed as session. closing

When an instance is under maintenance or the platform is about to close the connection, the Client may receive this event first, and then the connection will be closed. It is recommended to reconnect after receiving this event to avoid server-side interruption


2.4 Audio Formats and Stream Pushing

Input Audio

  • Format: PCM / Opus (Opus is automatically converted to PCM by the server), mono, 16000Hz, int16, little-endian;

  • It is recommended that audio sharding be in units of 20 ms (640 bytes per packet in 16k/int16 format), and be sent strictly in accordance with the rhythm of the real-time audio stream. Any deviation of the sending rate from the actual rhythm (either too fast or too slow) will trigger a server-side error;

  • The audio bytes are encoded in Base64 and placed in the audiofield.

Output Audio

  • Default: OGG-Opus; can be configured in session's audio.output.formator extension.tts.audio_configas PCM (24000Hz, mono, 16-bit little-endian);

  • The streaming block is located in the response.output_audio.delta deltafield (Base64).

2.5 Common Error Codes

Error Code Keywords Description and Handling Recommendations
42000020 volc_websearch_bot_id is required etc. web_agent networking mode missing bot_id/ missing api_key (check configuration)
45000003 Abnormal silence audio If there is no interaction for more than 10 minutes, the server will release the connection
50000000 AudioQueryError Model inference error
— found unknown escape character speaking_style/ system_rolecontains invalid characters. It is recommended to check the prompt.
55000001 ServerError Model inference error
55000001 ContextCanceled Not sent properly session. closemeans disconnection; be sure to disconnect only after receiving a reply
55000001 ClientError:InvalidSpeaker Invalid tone name
55000001 ExceededConcurrentDurationLimit Concurrent duration exceeds limit
50700000 stream recv timeout Model inference timed out
52000022 AudioChatError Model inference error
52000035 S2SQueryConnectError Model inference error

2.6 Notes

  1. Graceful shutdown: First call session. close, then close the WebSocket after receiving session. closed.

  2. Silent keepalive on the model side: send input_audio_mute. commitafter the microphone is turned off, and send input_audio_unmute. commitwhen resuming.

  3. Keep alive: input_audio_mute. commit only guarantees maintaining a session with the model side, not a connection with the platform. It is recommended to send a mute packet every 10 seconds to maintain the connection with the platform; the platform will release the connection if no mute packet is received for more than 90 seconds. If there is no interaction for more than 10 minutes, the model side may release the connection (error code 45000003).

  4. event_id: It is recommended that all upstream events carry the event_id.

  5. Context : conversation.item.createis initialized with user/assistant pairs; the timestamp is either full (incremented and not exceeding the current time) or not at all.

  6. Function Calling: session. updateperforms full coverage on tools; tool results are posted back in pairs by call_id; parallel calls are executed separately and then aggregated for a one-time postback.

  7. extension: The proprietary capabilities (ASR/TTS/Dialog) are uniformly placed in session. extensionand delivered along with session. create/ session. update.

  8. In scenarios such as instance maintenance, the Client may first receive the session. closing (see §2.3.21) delivered by the platform, and then the connection will be closed. It is recommended to reconnect after receiving this event to avoid server-side interruption.