Skip to content
← Tags

#Video Editing (1)

Video GenerationMediaTopA

Gemini Omni: moving the unit of video delivery from one attempt to a conversational editing process

Google DeepMind's video and multimodal generation model, officially create anything from any input, starting with video: it edits video through step-by-step conversation, each edit building on the last and keeping the scene coherent, raising the unit of delivery from one shot to footage you can keep revising; the other two surfaces are real-world knowledge and arbitrary reference composition. Read the level with its votes: on the Artificial Analysis arena this site syncs (2026-09-22) gemini-omni-1.1-flash is #1 for text-to-video at Elo 1516 with only 1,784 votes and a ±15 CI, while #2 is the sibling at 1513 with 26,576 votes and ±9, a 3-point gap far smaller than either CI, so statistically indistinguishable: the honest statement is that the Omni family ties for the top. For image-to-video minimax-h3 leads (1495, 57,112 votes) with Omni #2 at 1488 on 3,734 votes. Pricing: $1.50 per million input tokens, $9.00 text and $17.50 video output, 720p about $0.10 per second and a 30-second clip roughly $3. GA to paid-tier developers, none on the free tier. Closed, not self-hostable, not fine-tunable. Graded A (confirmed), our first A-grade video asset; A means the reading is credible and checkable, not undisputed first place.

1516Arena T2V Elo(1784 票)Confirmed · 2026-09
ProductGoogle DeepMindSite
Gemini Omni: moving the unit of video delivery from one attempt to a conversational editing process