SYSTEM Cited by 1 source
LiteLLM¶
Definition¶
LiteLLM is an open-source proxy / SDK that presents a unified API over many LLM providers (OpenAI, AWS Bedrock, Google Vertex, etc.), so client code targets one endpoint regardless of the underlying model. It is the substrate of the Zalando LLM proxy. (Source: sources/2026-08-13-zalando-agentic-engineering-at-zalando-a-snapshot)
Why Zalando chose it¶
Zalando's ML-platform team likes LiteLLM for its extensibility:
- Pre-call hooks — Zalando uses them to enforce client-version upgrades by User-Agent restriction.
- Post-call hooks — used for anonymized cost tracking.
- Prompt-caching auto-injection — checkpoints injected on behalf of custom-agent authors (concepts/context-engineering).
Operational caveats (from Zalando)¶
- Memory-leak / stability issues are real at scale: Zalando enforces a
restart every 20k requests via
--max_requests_before_restart. Even so, they run 2k MAU on six small (2 vCPU / 4 GB) pods. - Zalando "look[s] forward to the Rust rewrite that's expected to improve performance and stability."