Skip to content

SYSTEM Cited by 1 source

Gemini

Gemini is Google DeepMind's family of multimodal foundation models (text, image, video, audio) that serves as the base LLM / MLLM / VLM substrate across Google's product and research stack. The wiki tracks specific serving variants (e.g. Gemini 3.5 Flash) and tooling (e.g. Gemini CLI); this page is the canonical proper-noun anchor for the model family.

Role as an agent-orchestration substrate

In the long-form-video co-director work, Gemini is the foundation model that a multi-agent orchestration layer steers. (Source: sources/2026-09-24-google-automating-coherent-long-form-video-generation)

  • Co-Director feeds structured, MAB-selected creative prompts directly into Gemini (and Veo), while remaining model-agnostic — it can sit on top of any foundation generative model.
  • CANVAS and A²RD are built as orchestration layers on top of Gemini, adding persistent visual memory and agentic autoregressive generation respectively.
  • Gemini-3.1-Pro is used as the direct-generation baseline in the museum- heist continuity evaluation (it exhibits prop inconsistency + background drift without the CANVAS memory layer), and Gemini 3 Pro is paired with image models to construct the GenAD-Bench benchmark.
  • A multimodal Gemini acts as the MLLM Judge (concepts/llm-as-judge) that critiques the compiled cut and feeds a factored reward back into Co-Director's bandit, and as the VLM producing "semantic gradients" in VQQA.

Native safety

Content generated through Gemini (and Veo) inherently carries SynthID watermarking, which any orchestration layer built on top of it inherits for free.

Seen in

Last updated · 766 distilled / 2,225 read