Skip to content

SYSTEM Cited by 1 source

Zalando LLM proxy

Definition

Zalando's ML-platform team's central API proxy for LLM access, deployed January 2024 on top of LiteLLM. It gives every engineer API-based access to models from multiple providers (OpenAI, AWS Bedrock, Google Vertex) behind one endpoint, and gives the platform team a single chokepoint for measurement, cost control, and policy enforcement. Canonical Zalando instance of the LLM access proxy concept and the LLM-proxy-gateway pattern. (Source: sources/2026-08-13-zalando-agentic-engineering-at-zalando-a-snapshot)

What it does

  • Measurement: single point to track adoption — MAU, WAU, model, and User-Agent.
  • Pre-call hooks: enforce client-version upgrades by restricting proxy access based on the User-Agent header (for self-managed client installs, blocking is "the only effective measure"; same for retiring models).
  • Post-call hooks: anonymized cost tracking.
  • Prompt-caching auto-injection: injects prompt-caching checkpoints so custom-agent authors get the cost saving before they understand caching.
  • Vendor independence: onboarding a new provider is a proxy change; users keep their preferred tool.

Operational numbers

  • Serves 2k MAU on just six small pods (2 vCPU / 4 GB each).
  • Forced restart after 20k requests (--max_requests_before_restart) to mitigate LiteLLM memory-leak / stability issues; Zalando notes it looks forward to LiteLLM's Rust rewrite for performance/stability.

Companion surfaces

  • A simple chat UI (fork of an now-unmaintained OSS codebase) with surprisingly high adoption despite ubiquitous IDE/CLI alternatives.
  • A CLI tool built with pydantic-ai (incepted at an Aug 2024 hackathon before coding agents existed) that grew a maintainer community and features: image generation, interactive multi-turn chat with load/save-file context management, agent mode with MCP support + automatic Bearer-token injection for internal MCP servers, an http↔stdio MCP proxy, built-in MCP server configuration, and a coding agent configuration command that installs safe configs for claude code, opencode, and pi with model-autodiscovery plugins.
  • A local auth-injecting proxy that injects auth headers (bridging tools that only support static credentials / subscription auth) and ships a TUI showing per-model cost, prompt-cache-usage gaps, and per-request metadata (User-Agent, model, cost, token stats incl. cache read/write).
Last updated · 766 distilled / 2,225 read