Updated
Updated · Google Research · Sep 24
Google Unveils AI Video Co-Director, Generating 10-Minute Narratives With 81.4 Benchmark Score
Updated
Updated · Google Research · Sep 24

Google Unveils AI Video Co-Director, Generating 10-Minute Narratives With 81.4 Benchmark Score

3 articles · Updated · Google Research · Sep 24

Summary

  • Google said its new multi-agent AI video co-director can generate minutes-long, temporally consistent narratives, aiming to reduce character drift, scene changes and error propagation that plague chained video pipelines.
  • Built as an orchestration layer on top of Gemini and Veo, the system uses an orchestrator, storyboard and production agents, then a multimodal judge that feeds reward signals back into iterative refinement loops.
  • Google paired the core framework with CANVAS for persistent visual memory, A²RD for segment-by-segment long-video generation, and VQQA for prompt-based artifact correction without editing model internals.
  • In evaluations, the co-director reached a peak 81.4 score on GenAD-Bench, while Google said the broader toolkit improved continuity, long-duration consistency and compositional quality across several video benchmarks.
  • The company framed the research as a creator aid rather than a replacement, saying future work will add more human-in-the-loop controls for long-horizon visual storytelling.

Insights

Will Google's new AI co-director actually empower filmmakers, or will its automated optimizations strip away the unpredictable magic of human creativity?
With Veo 3.1 tracking every character's state, how long until these hyper-coherent AI simulations replace traditional film sets entirely?
Can persistent visual memory finally cure AI video's hallucination problem, or will it just create more expensive, complex rendering bottlenecks?