SYSTEM Cited by 1 source
Co-Director¶
Co-Director (the "AI video co-director") is Google Research's hierarchical multi-agent orchestration framework that formalizes multi-shot video storytelling as a global optimization problem. It is the top-level framework in Google's long-form-video suite (to appear at COLM 2026) and sits as an orchestration layer on top of Gemini and Veo — though its architecture is model-agnostic (it can steer any foundation generative model). (Source: sources/2026-09-24-google-automating-coherent-long-form-video-generation)
Problem it solves¶
Linear, handcrafted prompt chains produce semantic drift, content collapse, and cascading failures (an upstream artifact corrupts downstream synthesis). Because errors propagate and terminal failures can't be traced to a specific prompt, this is the classical credit-assignment problem — the same diagnosis behind cascading failure. Co-Director replaces rigid chains with hierarchical parameterization and a global optimization loop.
Architecture¶
Two interconnected loops: strategic steering + multi-stage production.
- Orchestrator Agent + multi-armed bandit. A multi-armed bandit (MAB) globally selects a creative configuration across three dimensions:
- Creative Strategy (intent),
- Narrative Mode (story structure),
- Aesthetic Archetype (visual tone & cinematography).
This casts creation as a search balancing exploration of novel narrative strategies against exploitation of effective configurations. The chosen config is injected into sub-agents' system prompts as top-down steering, guaranteeing the whole pipeline runs under one unified vision.
- Production hierarchy (driven by the config):
- Pre-Production Agent — synthesizes a scene-by-scene storyline + visual assets into a unified storyboard.
- Production Agent — translates the storyboard into audiovisual media via
specialized sub-agents:
- Keyframe Agent — anchors character & scene visuals,
- Video Agent — adds motion,
- Audio Agent — layers voiceover + score.
This is the specialized-agent- decomposition shape (each agent carries a small, focused scope).
- MLLM Judge → factored reward → MAB. A multimodal LLM (LLM-as-judge) critiques the compiled cut across the three designated dimensions and feeds a factored reward signal back to the MAB, which iteratively refines choices across successive generation loops. Quality is modeled as a test-time objective, decoupling creative synthesis from consistency.
Safety¶
As an orchestration layer feeding structured prompts into Gemini/Veo, all generated image/video/audio inherently carry SynthID watermarking; for production, additional safety classifiers can be applied over the final video to guard against unintended contextual interactions between individually-safe clips.
Results¶
- Peak quality score 81.4 on GenAD-Bench (50 fictional brands × 4 products = 400 scenarios, built via a human-in-the-loop pipeline pairing Gemini 3 Pro with image models).
- Enhanced story consistency on ViStoryBench.
Seen in¶
- sources/2026-09-24-google-automating-coherent-long-form-video-generation — the introducing post; Co-Director is the orchestration pillar of the suite.
Related¶
- systems/canvas, systems/a2rd, systems/vqqa — sibling frameworks in the same suite.
- systems/gemini, systems/veo, systems/synthid — the foundation-model + safety substrate.
- patterns/specialized-agent-decomposition, concepts/llm-as-judge, concepts/cascading-failure.
- companies/google.