SYSTEM Cited by 1 source
Zalando LLM proxy¶
Definition¶
Zalando's ML-platform team's central API proxy for LLM access, deployed January 2024 on top of LiteLLM. It gives every engineer API-based access to models from multiple providers (OpenAI, AWS Bedrock, Google Vertex) behind one endpoint, and gives the platform team a single chokepoint for measurement, cost control, and policy enforcement. Canonical Zalando instance of the LLM access proxy concept and the LLM-proxy-gateway pattern. (Source: sources/2026-08-13-zalando-agentic-engineering-at-zalando-a-snapshot)
What it does¶
- Measurement: single point to track adoption — MAU, WAU, model, and User-Agent.
- Pre-call hooks: enforce client-version upgrades by restricting proxy access based on the User-Agent header (for self-managed client installs, blocking is "the only effective measure"; same for retiring models).
- Post-call hooks: anonymized cost tracking.
- Prompt-caching auto-injection: injects prompt-caching checkpoints so custom-agent authors get the cost saving before they understand caching.
- Vendor independence: onboarding a new provider is a proxy change; users keep their preferred tool.
Operational numbers¶
- Serves 2k MAU on just six small pods (2 vCPU / 4 GB each).
- Forced restart after 20k requests (
--max_requests_before_restart) to mitigate LiteLLM memory-leak / stability issues; Zalando notes it looks forward to LiteLLM's Rust rewrite for performance/stability.
Companion surfaces¶
- A simple chat UI (fork of an now-unmaintained OSS codebase) with surprisingly high adoption despite ubiquitous IDE/CLI alternatives.
- A CLI tool built with pydantic-ai (incepted
at an Aug 2024 hackathon before coding agents existed) that grew a
maintainer community and features: image generation, interactive
multi-turn chat with load/save-file context management, agent mode with
MCP support + automatic Bearer-token injection for internal MCP servers,
an http↔stdio MCP proxy, built-in MCP server configuration, and a
coding agent configurationcommand that installs safe configs for claude code, opencode, and pi with model-autodiscovery plugins. - A local auth-injecting proxy that injects auth headers (bridging tools that only support static credentials / subscription auth) and ships a TUI showing per-model cost, prompt-cache-usage gaps, and per-request metadata (User-Agent, model, cost, token stats incl. cache read/write).
Related¶
- systems/litellm · systems/amazon-bedrock · systems/google-vertex-ai · systems/openai-api · systems/pydantic-ai · systems/opencode · systems/pi-coding-agent
- llm-access-proxy · user-agent-client-attribution · concepts/context-engineering · vendor-independence-for-llm-tooling
- patterns/ai-gateway-provider-abstraction · auth-injecting-local-proxy
- companies/zalando