SYSTEM Cited by 1 source
VQQA¶
VQQA (Video Quality Question Answering) is Google Research's unified, multi-agent framework for closed-loop refinement — letting the system autonomously identify and fix visual artifacts across diverse input modalities and video-generation tasks. (Source: sources/2026-09-24-google-automating-coherent-long-form-video-generation)
Core mechanism: VLM critiques as "semantic gradients"¶
VQQA dynamically generates visual questions tailored to the specific prompt, then uses the resulting Vision-Language Model (VLM) critiques as semantic gradients — natural-language directional feedback that guides iterative refinement, analogous to numerical gradients in backpropagation. This replaces passive evaluation metrics with human-interpretable, actionable feedback.
The refinement loop (over a natural-language interface):
generate video → evaluate via visual questions → refine the text prompt → repeat
Black-box prompt optimizer, not a pixel editor¶
In practice VQQA operates as a black-box prompt optimizer rather than a pixel-level editor:
- It does not mask or paint over frames; it iteratively refines the text prompt to correct high-level compositional defects (attribute-binding errors, inconsistent character attributes).
- The updated prompt guides the generator to sample a new path in latent space — e.g. rendering a realistic mylar-balloon texture onto a strict cuboid geometry, or fixing a mid-performance instrument change by keeping the violinist and pianist anchored to their instruments across cuts.
This matters architecturally: existing test-time-optimization methods are either computationally expensive or require white-box access to model internals. VQQA needs neither — it is the prompt- optimizer flywheel applied at inference/refinement time, and a closed-loop detect→critique→act cycle over generated media.
Global Selection: a drift-guard on the refinement loop¶
To prevent semantic drift during refinement, VQQA employs a Global Selection mechanism: rather than blindly taking the final iteration's output, a global VLM rater evaluates every video generated across the optimization trajectory against the original, unedited prompt, then selects the highest-scoring candidate. This ensures localized corrections do not compromise the broader context — a two-stage local-refine-then-global-select guard (cf. patterns/two-stage-evaluation).
Results¶
Notable absolute quality gains across T2V-CompBench, VBench2, and VBench-I2V, resolving physical and compositional inconsistencies.
Seen in¶
- sources/2026-09-24-google-automating-coherent-long-form-video-generation — the closed-loop-refinement pillar of the suite.
Related¶
- systems/co-director, systems/canvas, systems/a2rd — sibling frameworks.
- systems/gemini — the underlying VLM/foundation model.
- patterns/prompt-optimizer-flywheel, patterns/closed-loop-remediation, patterns/two-stage-evaluation, concepts/llm-as-judge.
- companies/google.