curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "google/gemini-3.1-flash/text-to-speech",
"input": {
"text": "Welcome to the studio. [short pause] Let us tell your story.",
"style_instructions": "Warm narration with a relaxed pace.",
"voice": "Kore",
"temperature": 1,
"output_format": "mp3"
}
}
'{
"code": 200,
"data": {
"task_id": "task-unified-example",
"status": "running",
"created_time": "2026-09-10T08:00:00"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}Google
Gemini 3.1 Flash TTS
Generate speech with Gemini 3.1 Flash TTS, using style instructions and selectable voices for narration or dialogue with two speakers.
POST
/
api
/
generate
/
submit
curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "google/gemini-3.1-flash/text-to-speech",
"input": {
"text": "Welcome to the studio. [short pause] Let us tell your story.",
"style_instructions": "Warm narration with a relaxed pace.",
"voice": "Kore",
"temperature": 1,
"output_format": "mp3"
}
}
'{
"code": 200,
"data": {
"task_id": "task-unified-example",
"status": "running",
"created_time": "2026-09-10T08:00:00"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}- Submit the request with
VIDGO_API_KEY. Atask_idis returned immediately. Keep it until the task reachesfinishedorfailed. - Use Query Task Status to retrieve the result. If you provide a
callback_url, Vidgo sends the result to your webhook when the task finishes or fails.
Gemini 3.1 Flash TTS
Generate speech with Gemini 3.1 Flash TTS, using style instructions and selectable voices for narration or dialogue with two speakers.Available Models
- google/gemini-3.1-flash/text-to-speech - Convert text into speech.
Output Options
Output Format
output_format: Optional. Accepted values:mp3,wav,ogg_opus. Default:mp3.
Key Features
- Convert text into speech.
- Use one voice or two speaker aliases.
- Control delivery with
style_instructionsandtemperature. - Select a voice with
voiceand specify the spoken language withlanguage_code.
Advanced Parameters
Setmodel and optional callback_url at the request root. Place all workflow parameters, including output options, inside input.
Use the fields in the parameter reference to configure speech generation.
- When supplying
speakers, provide exactly two distinctspeaker_idaliases using only letters, digits, and underscores. Speech uses the voice assigned to each speaker.
Text
text: Required. Supply the text to speak. Length: 1-50,000 characters.- Include the words to speak, with optional audio tags such as
[short pause],[whispering], and[laughing]. For dialogue, prefix each turn with its matchingspeaker_idfollowed by a colon.
Style Instructions
style_instructions: Optional. Describe how the speech should be delivered.
Voice
voice: Optional. Select the voice. String. Default:Kore.
Supported voice values
Supported voice values
AchernarAchirdAlgenibAlgiebaAlnilamAoedeAutonoeCallirrhoeCharonDespinaEnceladusErinomeFenrirGacruxIapetusKoreLaomedeiaLedaOrusPulcherrimaPuckRasalgethiSadachbiaSadaltagerSchedarSulafatUmbrielVindemiatrixZephyrZubenelgenubi
Language Code
language_code: Optional. Select a value below, or omit it to detect the language fromtext.
Supported language code values
Supported language code values
Arabic (Egypt)Bangla (Bangladesh)Dutch (Netherlands)English (India)English (US)French (France)German (Germany)Hindi (India)Indonesian (Indonesia)Italian (Italy)Japanese (Japan)Korean (South Korea)Marathi (India)Polish (Poland)Portuguese (Brazil)Romanian (Romania)Russian (Russia)Spanish (Spain)Tamil (India)Telugu (India)Thai (Thailand)Turkish (Turkey)Ukrainian (Ukraine)Vietnamese (Vietnam)Afrikaans (South Africa)Albanian (Albania)Amharic (Ethiopia)Arabic (World)Armenian (Armenia)Azerbaijani (Azerbaijan)Basque (Spain)Belarusian (Belarus)Bulgarian (Bulgaria)Burmese (Myanmar)Catalan (Spain)Cebuano (Philippines)Chinese Mandarin (China)Chinese Mandarin (Taiwan)Croatian (Croatia)Czech (Czech Republic)Danish (Denmark)English (Australia)English (UK)Estonian (Estonia)Filipino (Philippines)Finnish (Finland)French (Canada)Galician (Spain)Georgian (Georgia)Greek (Greece)Gujarati (India)Haitian Creole (Haiti)Hebrew (Israel)Hungarian (Hungary)Icelandic (Iceland)Javanese (Java)Kannada (India)Konkani (India)Lao (Laos)Latin (Vatican City)Latvian (Latvia)Lithuanian (Lithuania)Luxembourgish (Luxembourg)Macedonian (North Macedonia)Maithili (India)Malagasy (Madagascar)Malay (Malaysia)Malayalam (India)Mongolian (Mongolia)Nepali (Nepal)Norwegian Bokmal (Norway)Norwegian Nynorsk (Norway)Odia (India)Pashto (Afghanistan)Persian (Iran)Portuguese (Portugal)Punjabi (India)Serbian (Serbia)Sindhi (India)Sinhala (Sri Lanka)Slovak (Slovakia)Slovenian (Slovenia)Spanish (Latin America)Spanish (Mexico)Swahili (Kenya)Swedish (Sweden)Urdu (Pakistan)
Speakers
-
speakers: Optional. Supply exactly 2 items. Each item is an object. -
speakers[].voice: Required. Select a voice for this speaker.
Supported voice values
Supported voice values
AchernarAchirdAlgenibAlgiebaAlnilamAoedeAutonoeCallirrhoeCharonDespinaEnceladusErinomeFenrirGacruxIapetusKoreLaomedeiaLedaOrusPulcherrimaPuckRasalgethiSadachbiaSadaltagerSchedarSulafatUmbrielVindemiatrixZephyrZubenelgenubi
speakers[].speaker_id: Required. String. Use only letters, digits, and underscores.
Temperature
temperature: Optional. Control speech-generation variability. Number from0to2. Default:1.
Output Files
- Read every item in the returned
filesarray. - Generated audio uses
file_type: "audio". - Read the audio URL from
files[].file_url.
Authorizations
Use VIDGO_API_KEY.
Body
application/json
Set model and optional callback_url at the request root. Place speech parameters inside input.