Skip to content

SYSTEM Cited by 1 source

A²RD

A²RD is Google Research's agentic autoregressive video generation architecture — the pillar that translates storyboards into actual minutes-long video with consistent narrative progression across multi-minute temporal gaps. (Source: sources/2026-09-24-google-automating-coherent-long-form-video-generation)

Core mechanism: segment-by-segment generation with multimodal video memory

A²RD generates video segment-by-segment, augmented with a multimodal video memory that tracks segment contexts and dynamics. For each segment it runs a four-step loop:

retrieve → synthesize → refine → update

The memory substrate is the anti-drift mechanism — it is agent memory specialized to video world-state, continuously queried to maintain character identity, costume details, and structural geometry from the opening shot to the final frame (standard video generators instead suffer visual decay — characters mutate, locations morph).

Adaptive mode switching: extrapolation vs interpolation

A critical part of the loop is that the agent adaptively determines the segment generation mode, smoothly switching between:

  • Extrapolation — allow natural narrative progression (push the plot into new beats), and
  • Interpolation — anchor segments to existing entities and environments (preserve the physical reality of the scene).

This balances moving the story forward against maintaining continuity — the long-horizon-temporal analog of the world-state tracking CANVAS does at the storyboard layer.

Results

Minimizes layout drift and improves character/environment consistency over continuous multi-minute runs on VBench-Long and LVBench-C (120 text-only scenarios with a strict ≥10-segment disappearance gap rule before assets return with narrative-driven changes).

Seen in

Last updated · 766 distilled / 2,225 read