curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "kwaivgi/kling-video-o3-4k/text-to-video",
"input": {
"duration": 4,
"sound": true,
"multi_shots": false,
"prompt": "Photoreal slow lateral dolly through an empty ancient temple gallery in rain. Weathered floral stone reliefs fill the foreground; rows of pillars recede toward a quiet courtyard. Water follows carved grooves and drips from worn edges, revealing mineral grains, chisel marks and moss in cracks. Soft overcast daylight, subtle wet highlights, stable architecture and rich fine detail. Audio: gentle rain and isolated drips; no music. No text, logos or watermarks.",
"aspect_ratio": "16:9"
}
}
'{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}Kling O3 4K Text To Video
Generate 3–15-second videos from text prompts with Kling O3 4K. Outputs 4K video. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
curl --request POST \
--url https://api.vidgo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "kwaivgi/kling-video-o3-4k/text-to-video",
"input": {
"duration": 4,
"sound": true,
"multi_shots": false,
"prompt": "Photoreal slow lateral dolly through an empty ancient temple gallery in rain. Weathered floral stone reliefs fill the foreground; rows of pillars recede toward a quiet courtyard. Water follows carved grooves and drips from worn edges, revealing mineral grains, chisel marks and moss in cracks. Soft overcast daylight, subtle wet highlights, stable architecture and rich fine detail. Audio: gentle rain and isolated drips; no music. No text, logos or watermarks.",
"aspect_ratio": "16:9"
}
}
'{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}"Request timeout"{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}{
"detail": "<string>"
}Parameters
Use modelkwaivgi/kling-video-o3-4k/text-to-video. Put generation parameters inside input.
prompt: Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true.multi_prompt: Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.duration: Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value.multi_shots: Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error.sound: Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true.aspect_ratio: Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio.
Pricing
50 credits/s with or without sound. Credits = duration × rate. Failed generation tasks are refunded.Default example: Rain Along the Stone Gallery
This is the same default example shown on the model page.{
"model": "kwaivgi/kling-video-o3-4k/text-to-video",
"input": {
"duration": 4,
"sound": true,
"multi_shots": false,
"prompt": "Photoreal slow lateral dolly through an empty ancient temple gallery in rain. Weathered floral stone reliefs fill the foreground; rows of pillars recede toward a quiet courtyard. Water follows carved grooves and drips from worn edges, revealing mineral grains, chisel marks and moss in cracks. Soft overcast daylight, subtle wet highlights, stable architecture and rich fine detail. Audio: gentle rain and isolated drips; no music. No text, logos or watermarks.",
"aspect_ratio": "16:9"
}
}
Completed result from the status endpoint
GET /api/generate/status/{task_id}
The response below records this verified example. For a new generation, query the task ID returned by your own submission.
{
"code": 200,
"data": {
"task_id": "EE4XVIYVL5IAX4YG",
"status": "finished",
"files": [
{
"file_type": "video",
"file_url": "https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-4k/text-to-video/v1/01/output.mp4"
}
],
"created_time": "2026-09-22T18:48:45",
"error_message": null,
"progress": 100
}
}
Authorizations
Use VIDGO_API_KEY.
Body
Keep model and callback_url at the root. Unknown root fields are ignored; unsupported input fields are rejected.
Vidgo public model ID. Must be kwaivgi/kling-video-o3-4k/text-to-video.
kwaivgi/kling-video-o3-4k/text-to-video "kwaivgi/kling-video-o3-4k/text-to-video"
Video generation parameters. Use standard JSON numbers and booleans. For compatibility, integer strings such as "5" are accepted. Boolean strings true/1/yes/y/on mean true; false/0/no/n/off mean false. These strings are case-insensitive and trimmed. Numeric 1 and 0 are also accepted for boolean fields. Numeric and boolean prompt values are converted to text; 0, false and null are treated as empty. Objects and arrays are not accepted as prompts.
- Option 1
- Option 2
Show child attributes
Show child attributes
Optional public HTTP(S) endpoint for task completion notifications. HTTPS is recommended. Omit, use null, or use an empty string to disable callbacks.