Skip to content

Agentic Engineering at Zalando: a snapshot

Summary

A survey/retrospective from Zalando Engineering (2026-08-13) looking back on ~2.5 years of adopting LLMs and agentic engineering across 250+ engineering teams. Rather than a single-system deep dive, it documents the platform substrate Zalando built to let a large, decentralized org experiment safely: a LiteLLM-based LLM access proxy shipped in January 2024, a companion chat UI + pydantic-ai CLI, an Identity Broker for OAuth2 on-behalf-of delegation to MCP servers, a risk-based PR-approval bot, a centralized agent-skill collection, and structured knowledge-sharing formats (LLM guild, GenAI Labs, hackathons). The architectural through-line is vendor independence: a single proxy chokepoint gives measurement, cost control, prompt-caching auto-injection, and client-version enforcement while never mandating a single tool.

Key takeaways

  1. A single LLM proxy is the load-bearing platform primitive. The ML platform team deployed a LiteLLM-based API proxy in Jan 2024 fronting OpenAI, AWS Bedrock, and Google Vertex. It became the single point to measure adoption (MAU, WAU, model, User-Agent) and to enforce policy. (Source: sources/2026-08-13-zalando-agentic-engineering-at-zalando-a-snapshot)
  2. Operational numbers on the proxy: runs 2k MAU on just six small pods (2 vCPU / 4 GB each); enforces a restart after 20k requests (--max_requests_before_restart) to mitigate LiteLLM memory leaks / instability. Uses pre-call hooks to enforce client-version upgrades by User-Agent header and post-call hooks for anonymized cost tracking. Auto-injects prompt-caching checkpoints so custom-agent authors get the cost saving before they learn about caching.
  3. Vendor independence is an explicit design goal, and Zalando never centrally mandated a single tool. Users pick IDE vs CLI, model provider, and coding agent by preference; opencode and pi are called out as a sweet-spot because they mix a GitHub Copilot subscription with the API proxy. Observation: switching costs are low but psychological attachment to a tool/model is real, so nudging (limits/errors) matters more than mandates.
  4. Risk-based PR approval cut PR lead time 20–40%. A bot scores each PR at creation as low / medium / high rollout risk; 33% of PRs are low-risk and auto-approved (risk-based-auto-approval). The ruleset is derived from analysis of Zalando's own production incidents and is highly specific to their stack (config typos = high risk — would have caught the metadpata incident; backwards-incompatible change = medium; docs-only = low). Second-order effect: engineers restructure PRs to be low-risk (split backwards-compatible changes from field-removals).
  5. The auth problem for agentic tooling is the hard part, and Zalando is solving it centrally. A local auth-injecting proxy + an http↔stdio MCP proxy inject Bearer tokens for internally-hosted MCP servers so no secrets are hardcoded in config files; deployed MCP servers are auto-protected by a default ingress OAuth filter. The forthcoming Identity Broker captures delegation chains for on-behalf-of flows, brokers between OAuth2 infrastructures, and implements a token vault — sitting in the call path (via an infra gateway) between agent↔MCP-server and agent↔agent.
  6. AI coding is empirically visible in the codebase and in process data. PR sizes grew (esp. the [500,1k) and [1k,2k) buckets after Sonnet 4 in Q2/2025); commit-message size clusters ~5k chars (one included a full unit-test log); total cyclomatic complexity shows inflection points that line up with agent adoption (AI amplifies good and bad practices). Agent-only greenfield services build complexity fast then plateau.
  7. Session data is highly educational. Inspecting coding-agent sessions (via agentsview / codeburn) surfaces wasteful traffic (plan names, terminal titles, idle recaps) and prompting anti-patterns; a homemade parser found one user's opencode cache-hit ratio at <30% vs 80%+ expected.
  8. Governance stance: it's too early to converge. With 200+ teams exploring, the objective is transparency, not standardization — the internal Tech Radar grew an AI section (a departure from a decade of pushing library choices to language communities), surfaced through the Backstage-based Sunrise developer portal. AI model usage is auto-detected by scanning deployed Docker images and auto-registered for legal review.
  9. Skills, knowledge-sharing, and what's next. A centralized agent-skill collection (grouped into plugins, distributed via managed config or a symlink-installing CLI command) disseminates best practices — migration skills are the most popular. The agent platform being built composes OSS like kagent for Kubernetes agent runtime, plus the Identity Broker for auth; open problems named: device/config management, local sandboxing, and auto-routing across models (incl. open-weight).

Systems / concepts / patterns extracted

Systems: Zalando LLM proxy · LiteLLM · risk-based PR-approval bot · Identity Broker · kagent · pydantic-ai · opencode · pi · GitHub Copilot · AWS Bedrock · Google Vertex AI · OpenAI API · Backstage

Concepts: llm-access-proxy · user-agent-client-attribution · concepts/context-engineering · vendor-independence-for-llm-tooling · risk-based-pr-approval · on-behalf-of-flow · concepts/observability · agentic-code-complexity-amplification · agent-skill · concepts/token-vault · tech-radar-language-governance

Patterns: patterns/ai-gateway-provider-abstraction · auth-injecting-local-proxy · risk-based-auto-approval · patterns/on-behalf-of-agent-authorization · centralized-agent-skill-collection

Operational numbers

  • 250+ engineering teams; proxy serves 2k MAU on 6 pods (2 vCPU / 4 GB each).
  • LiteLLM restart forced every 20k requests to bound memory-leak risk.
  • Providers behind the proxy: OpenAI, AWS Bedrock, Google Vertex.
  • 33% of PRs auto-approved as low-risk → 20–40% PR lead-time reduction on those.
  • PR-size growth concentrated in [500,1k) and [1k,2k) buckets post-Sonnet-4 (Q2/2025); commit messages cluster near 5k chars.
  • One user's opencode cache-hit ratio <30% vs 80%+ expected (isolated, not systemic).
  • Codebases studied: go-agentic-only (new, full agent adoption), go-reference (10y+ OSS), java-with-agents (4y), java-reference (12y+, no agents).

Caveats

  • This is a snapshot / survey, not a production deep-dive: numbers are disclosed selectively and several components (Identity Broker, AI-readiness scanner, agent platform) are in-progress or forthcoming, not shipped at publication.
  • Complexity-inflection findings are correlational; agent-authorship markers (Co-authored-by) are inconsistent, especially in OSS, so attribution to agents is partial.
  • The risk-approval ruleset is highly Zalando-specific (their deployment manifests, config formats, incident history) — it's a methodology, not a portable ruleset.
  • Tool preferences and switching-cost claims are qualitative observations across the org, not a controlled study.

Source

Last updated · 766 distilled / 2,225 read