curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "kwaivgi/kling-video-o3-pro/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Cinematic whimsical realism in a warm oak library. Medium two-shot: a small ivory robot with an oval face and amber eyes shelves a blue book, accidentally knocking one red book onto the floor. Beside it, an adult librarian wears a green cardigan and round glasses. Both remain visible under warm reading lamps. One distinct book thud against quiet room tone. Unmarked book covers. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut closer to the same ivory robot and green-cardigan librarian in the same oak library. The librarian raises one finger to her lips. The robot tilts its head apologetically and softly says exactly \"Sorry.\" in a gentle robotic voice, synchronized with its small mouth light. Preserve their appearance, positions and warm lighting. End in an embarrassed pause; no music. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
'{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}Kling O3 Pro Text To Video
Generate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "kwaivgi/kling-video-o3-pro/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Cinematic whimsical realism in a warm oak library. Medium two-shot: a small ivory robot with an oval face and amber eyes shelves a blue book, accidentally knocking one red book onto the floor. Beside it, an adult librarian wears a green cardigan and round glasses. Both remain visible under warm reading lamps. One distinct book thud against quiet room tone. Unmarked book covers. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut closer to the same ivory robot and green-cardigan librarian in the same oak library. The librarian raises one finger to her lips. The robot tilts its head apologetically and softly says exactly \"Sorry.\" in a gentle robotic voice, synchronized with its small mouth light. Preserve their appearance, positions and warm lighting. End in an embarrassed pause; no music. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
'{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}Parameters
Use modelkwaivgi/kling-video-o3-pro/text-to-video. Put generation parameters inside input.
prompt: Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true.multi_prompt: Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.duration: Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value.multi_shots: Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error.sound: Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true.aspect_ratio: Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio.
Pricing
13 credits/s without sound; 16 credits/s with sound. Credits = duration × rate. Failed generation tasks are refunded.Default example: A Quiet Apology
This is the same default example shown on the model page.{
"model": "kwaivgi/kling-video-o3-pro/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Cinematic whimsical realism in a warm oak library. Medium two-shot: a small ivory robot with an oval face and amber eyes shelves a blue book, accidentally knocking one red book onto the floor. Beside it, an adult librarian wears a green cardigan and round glasses. Both remain visible under warm reading lamps. One distinct book thud against quiet room tone. Unmarked book covers. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut closer to the same ivory robot and green-cardigan librarian in the same oak library. The librarian raises one finger to her lips. The robot tilts its head apologetically and softly says exactly \"Sorry.\" in a gentle robotic voice, synchronized with its small mouth light. Preserve their appearance, positions and warm lighting. End in an embarrassed pause; no music. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
Completed result from the status endpoint
GET /api/generate/status/{task_id}
The response below records this verified example. For a new generation, query the task ID returned by your own submission.
{
"code": 200,
"data": {
"task_id": "BRDSCRKN4Q0WH4T6",
"status": "finished",
"files": [
{
"file_type": "video",
"file_url": "https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-pro/text-to-video/v1/01/output.mp4"
}
],
"created_time": "2026-09-22T18:43:52",
"error_message": null,
"progress": 100
}
}
Authorizations
Use VIDGO_API_KEY.
Body
Keep model and callback_url at the root. Unknown root fields are ignored; unsupported input fields are rejected.
Vidgo public model ID. Must be kwaivgi/kling-video-o3-pro/text-to-video.
kwaivgi/kling-video-o3-pro/text-to-video "kwaivgi/kling-video-o3-pro/text-to-video"
Video generation parameters. Use standard JSON numbers and booleans. For compatibility, integer strings such as "5" are accepted. Boolean strings true/1/yes/y/on mean true; false/0/no/n/off mean false. These strings are case-insensitive and trimmed. Numeric 1 and 0 are also accepted for boolean fields. Numeric and boolean prompt values are converted to text; 0, false and null are treated as empty. Objects and arrays are not accepted as prompts.
- Option 1
- Option 2
Show child attributes
Show child attributes
Optional public HTTP(S) endpoint for task completion notifications. HTTPS is recommended. Omit, use null, or use an empty string to disable callbacks.