Skip to main content
POST
  • Use gemini-omni-flash-preview-official in the model field
  • Supports text-to-video, image/reference-to-video, and video editing
  • Supports up to 10 reference images or up to 3 input videos; images and videos cannot be mixed
  • Video duration is limited to 10 seconds; the Playground offers 4 / 6 / 8 / 10 second presets
  • Only 720p resolution is supported; both 16:9 and 9:16 aspect ratios are available
  • Async task API: submit a task, then query by task ID
This page documents gemini-omni-flash-preview-official. The platform routes it through the Vertex AI Interactions API with Service Account OAuth, not the Vertex AI Publisher Model predictLongRunning endpoint. It remains separate from the legacy gemini-omni-flash / gemini_omni_flash model and does not change that model’s routing or parameters.
Upload local images with the Upload Image API and local videos with the Upload Video API first, then pass the returned URLs. Do not pass base64 media data directly.

Authorizations

string
required
Use Bearer Token authentication:

Body

string
default:"gemini-omni-flash-preview-official"
required
Model name. Use gemini-omni-flash-preview-official.
string
required
Text prompt for video generation.
integer
default:"6"
Video duration in seconds. The maximum is 10; the Playground offers 4, 6, 8, and 10 second presets.
string
default:"16:9"
Video aspect ratio:
  • 16:9 landscape
  • 9:16 portrait
string
default:"720p"
Video resolution:
  • 720p default resolution
  • Only 720p is supported; Veo-specific parameters are not accepted
string[]
Optional reference image URL array with up to 10 items. Omit it for text-to-video. This field cannot be combined with video_list.
object[]
Optional input videos for video editing, with up to 3 items. Each object must contain video_url, for example { "video_url": "https://example.com/input.mp4" }. HTTPS URLs, data URIs, and platform-uploaded video URLs are supported. Each downloaded input video is limited to 50 MB. This field cannot be combined with image_urls.
object
Optional Interactions generation controls. Supported keys include temperature (0.02.0), topP / top_p (0.01.0), candidateCount / candidate_count, previous_interaction_id, and task (text_to_video, image_to_video, reference_to_video, or edit). When task is omitted, the platform does not invent a default; the upstream model infers the mode from the prompt and media. For edit (or video inputs without an explicit non-edit task), do not rely on aspect_ratio; the platform omits it for that path because upstream rejects aspect ratio on edit tasks.

Response

string
Task ID for status polling.
string
Object type, usually generation.task.
string
Model name used for the request.
string
Task status: queued, in_progress, completed, or failed.
integer
Task creation timestamp.
string
Generated video URL when the task completes.

Examples

Video editing example:

Query Task

The generation endpoint returns a task ID. Use the common video task status endpoint to fetch status and results: