Skip to content

Google Research lays out four-agent system for coherent long-form AI video

The co-director framework orchestrates Gemini and Veo to track story worlds, refine prompts and limit drift across multiple shots.

Google Research has introduced an AI video co-director, a unified multi-agent framework intended to generate temporally consistent, long-form video narratives. The system combines Co-Director, CANVAS, A²RD and VQQA to automate multi-model prompting, shot chaining and visual refinement. The announcement, published Sept. 24 and authored by Google research scientists Yale Song and Yiwen Song, frames the work as an orchestration layer over Gemini and Veo. Its model-agnostic architecture can also operate with other generative foundation models. The aim is to address a central weakness of current video-generation workflows: individual diffusion-generated clips can be high fidelity, but lengthy narratives may suffer from changing costumes, characters, props and scenery. Google describes such pipelines as vulnerable to upstream artifacts that cascade into later synthesis, while early errors become hard to trace back to the prompt that caused them and may require substantial manual correction. Co-Director treats storytelling as a global optimisation and world-state tracking problem rather than a rigid chain of prompts. An Orchestrator Agent uses a multi-armed bandit to explore creative strategy, narrative mode and aesthetic archetype; a Pre-Production Agent then produces a scene-by-scene storyboard. A Production Agent coordinates keyframe, video and audio sub-agents, while a multimodal model judges the compiled result and returns a factored reward signal for further optimisation. CANVAS, short for Continuity-Aware Narratives via Visual Agentic Storyboarding, keeps structured representations of characters, locations and object states as the narrative advances. Its persistent visual memory retrieves anchors when a character or setting returns after a cutaway, seeking to preserve identity, spatial structure and object state across revisited scenes. In a museum-heist comparison, Google said Gemini-3.1-Pro alone changed props and room layouts, while AutoStudio lost character details across cuts. A²RD, or Agentic Autoregressive Diffusion, generates a video segment by segment. It uses multimodal video memory and a retrieve-synthesise-refine-update loop, alternating between extrapolation for new story beats and interpolation to anchor material to existing characters and environments. Google said it generated a 10-minute film, maintaining details such as character identity, costumes and structural geometry over long temporal gaps. VQQA, meaning Video Quality Question Answering, is designed to improve generation through prompt refinement rather than pixel editing. It creates prompt-specific visual questions, uses vision-language-model critiques as what Google calls “semantic gradients”, and applies a global selection mechanism that rates each candidate against the original prompt rather than simply accepting the final iteration. Google developed three specialised evaluation sets for the work. GenAD-Bench includes 400 advertising scenarios involving 200 fictional products from 50 brands; HardContinuityBench tests returning scenes and changes in prop state; and LVBench-C has 120 scenarios in which key assets disappear for at least 10 segments before reappearing. The company reported a peak quality score of 81.4 on GenAD-Bench, as well as consistency gains for the respective systems on several other video benchmarks. The framework inherits safeguards from its underlying models, including SynthID watermarking. SynthID is Google DeepMind technology for embedding digital watermarks in AI-generated media; for video, the watermark is embedded imperceptibly into every generated frame. The full four-part stack is not yet available as a downloadable product, public API or open-source implementation, according to AlphaSignal. Google said Co-Director is scheduled to appear at COLM 2026 and CANVAS at EMNLP 2026.