curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "elevenlabs/tts/eleven-v3",
"input": {
"text": "[whispers] Wait... that ticking is coming from inside the wall. [curious] This little brass key fits the crack. One turn, and... [excited] the whole ceiling is full of stars! [laughs] Grandpa, you built a planetarium in the attic. [exhales] All these years, it was waiting for us.",
"voice": "Charlotte",
"language_code": "en",
"stability": 0.5,
"timestamps": false,
"apply_text_normalization": "auto"
}
}
'
{
"code": 200,
"data": {
"task_id": "LPV36KJQ0Y2SYT21",
"status": "running",
"created_time": "2026-09-27T10:45:57"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}ElevenLabs
ElevenLabs V3 TTS
Turn text into spoken audio with ElevenLabs V3, with selectable voices, language control, and optional timestamps for narration and dialogue.
POST
/
api
/
generate
/
submit
curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "elevenlabs/tts/eleven-v3",
"input": {
"text": "[whispers] Wait... that ticking is coming from inside the wall. [curious] This little brass key fits the crack. One turn, and... [excited] the whole ceiling is full of stars! [laughs] Grandpa, you built a planetarium in the attic. [exhales] All these years, it was waiting for us.",
"voice": "Charlotte",
"language_code": "en",
"stability": 0.5,
"timestamps": false,
"apply_text_normalization": "auto"
}
}
'
{
"code": 200,
"data": {
"task_id": "LPV36KJQ0Y2SYT21",
"status": "running",
"created_time": "2026-09-27T10:45:57"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}- Submit the request with
VIDGO_API_KEY. Atask_idis returned immediately. Keep it until the task reachesfinishedorfailed. - Use Query Task Status to retrieve the result. If you provide a
callback_url, Vidgo sends the result to your webhook when the task finishes or fails.
ElevenLabs V3 TTS
Turn text into spoken audio with ElevenLabs V3, with selectable voices, language control, and optional timestamps for narration and dialogue.Available Models
- elevenlabs/tts/eleven-v3 - Convert text into speech.
Key Features
- Convert text into speech.
- Optionally request timestamps alongside the generated audio.
- Select a voice with
voiceand specify the spoken language withlanguage_code.
Advanced Parameters
Setmodel and optional callback_url at the request root. Place all workflow parameters, including output options, inside input.
Text
text: Required. Supply the text to speak. Length: 1-5,000 characters.- Leading and trailing whitespace is trimmed before the character limit is checked.
Voice
voice: Optional. Select the voice. String. Default:Rachel.- Available values:
Aria,Roger,Sarah,Laura,Charlie,George,Callum,River,Liam,Charlotte,Alice,Matilda,Will,Jessica,Eric,Chris,Brian,Daniel,Lily,Bill,Rachel.
Stability
stability: Optional. Control voice stability. Number from0to1. Default:0.5.
Timestamps
timestamps: Optional. Request timestamp information with the audio. Usetrueorfalse. Default:false.
Language Code
language_code: Optional. ISO 639-1 language code, such asenfor English orzhfor Chinese.
Apply Text Normalization
apply_text_normalization: Optional. Select text normalization behavior. Accepted values:auto,on,off. Default:auto.
Output Files
- Read every item in the returned
filesarray. - Generated audio uses
file_type: "audio". - Read returned
timestamps.jsonas a separate item withfile_type: "other". Preserve optional file metadata when present.
Examples
Each example pairs a request with its completed task response and generated audio. Submit the request toPOST https://api.vidgo.ai/api/generate/submit, then query GET https://api.vidgo.ai/api/generate/status/{task_id} with the task ID returned by your submission.
Stars in the attic
Stars in the attic
{
"model": "elevenlabs/tts/eleven-v3",
"input": {
"text": "[whispers] Wait... that ticking is coming from inside the wall. [curious] This little brass key fits the crack. One turn, and... [excited] the whole ceiling is full of stars! [laughs] Grandpa, you built a planetarium in the attic. [exhales] All these years, it was waiting for us.",
"voice": "Charlotte",
"language_code": "en",
"stability": 0.5,
"timestamps": false,
"apply_text_normalization": "auto"
}
}
{
"code": 200,
"data": {
"task_id": "LPV36KJQ0Y2SYT21",
"status": "finished",
"files": [
{
"file_url": "https://cdn.vidgo.ai/apis/models/elevenlabs/tts/eleven-v3/v1/01/speech.mp3",
"file_type": "audio",
"file_name": "speech.mp3"
}
],
"created_time": "2026-09-27T10:45:57",
"error_message": null,
"progress": 100
}
}
My first dumpling
My first dumpling
{
"model": "elevenlabs/tts/eleven-v3",
"input": {
"text": "[sighs] 我练了一个下午,面团还是歪歪扭扭的。奶奶看了一眼,说:别急,手放轻一点。[curious] 这样吗?先对折,再慢慢捏紧……[laughs] 真的站住了!虽然像只胖企鹅,但这是我第一次包好饺子。原来,有些本事不是听懂的,是一起做着学会的。",
"voice": "Rachel",
"language_code": "zh",
"stability": 0.5,
"timestamps": false,
"apply_text_normalization": "auto"
}
}
{
"code": 200,
"data": {
"task_id": "VLXQ5DI7PNWFRJQ4",
"status": "finished",
"files": [
{
"file_url": "https://cdn.vidgo.ai/apis/models/elevenlabs/tts/eleven-v3/v1/02/speech.mp3",
"file_type": "audio",
"file_name": "speech.mp3"
}
],
"created_time": "2026-09-27T10:46:38",
"error_message": null,
"progress": 100
}
}
The character of sound
The character of sound
{
"model": "elevenlabs/tts/eleven-v3",
"input": {
"text": "[curious] Why does a violin sound different from a flute when both play the same note? Their fundamental frequency can match, but the overtones do not. [excited] Those extra vibrations give each instrument its own sound. At 440 Hz, the pitch is A; the pattern above it is what makes the voice of the instrument recognizable.",
"voice": "George",
"language_code": "en",
"stability": 0.5,
"timestamps": true,
"apply_text_normalization": "on"
}
}
{
"code": 200,
"data": {
"task_id": "0E09515629QS87GR",
"status": "finished",
"files": [
{
"file_url": "https://cdn.vidgo.ai/apis/models/elevenlabs/tts/eleven-v3/v1/03/speech.mp3",
"file_type": "audio",
"file_name": "speech.mp3"
},
{
"file_url": "https://cdn.vidgo.ai/apis/models/elevenlabs/tts/eleven-v3/v1/03/timestamps.json",
"file_type": "other",
"file_name": "timestamps.json"
}
],
"created_time": "2026-09-27T10:47:13",
"error_message": null,
"progress": 100
}
}
Authorizations
Use VIDGO_API_KEY.
Body
application/json
Set model and optional callback_url at the request root. Place speech generation parameters inside input.