Google Unveils AI Video Co-Director, Generating 10-Minute Narratives With 81.4 Benchmark Score
Updated
Updated · Google Research · Sep 24
Google Unveils AI Video Co-Director, Generating 10-Minute Narratives With 81.4 Benchmark Score
3 articles · Updated · Google Research · Sep 24
Summary
Google said its new multi-agent AI video co-director can generate minutes-long, temporally consistent narratives, aiming to reduce character drift, scene changes and error propagation that plague chained video pipelines.
Built as an orchestration layer on top of Gemini and Veo, the system uses an orchestrator, storyboard and production agents, then a multimodal judge that feeds reward signals back into iterative refinement loops.
Google paired the core framework with CANVAS for persistent visual memory, A²RD for segment-by-segment long-video generation, and VQQA for prompt-based artifact correction without editing model internals.
In evaluations, the co-director reached a peak 81.4 score on GenAD-Bench, while Google said the broader toolkit improved continuity, long-duration consistency and compositional quality across several video benchmarks.
The company framed the research as a creator aid rather than a replacement, saying future work will add more human-in-the-loop controls for long-horizon visual storytelling.
Will Google's new AI co-director actually empower filmmakers, or will its automated optimizations strip away the unpredictable magic of human creativity?
With Veo 3.1 tracking every character's state, how long until these hyper-coherent AI simulations replace traditional film sets entirely?
Can persistent visual memory finally cure AI video's hallucination problem, or will it just create more expensive, complex rendering bottlenecks?