Skip to content

CLOUDFLARE

Source: The Cloudflare Blog — Brought to you by EmDash

Summary

On 2026-08-12 Cloudflare migrated its own blog (blog.cloudflare.com) off its previous CMS vendor onto EmDash, the open-source Astro-based CMS Cloudflare built, running on Cloudflare Workers. As "Customer Zero" for EmDash, Cloudflare validated the platform against real blog traffic — normal load ~75 RPS spiking above 5,000 RPS — using k6 load-test scenarios (ramp, breakpoint, burst) with explicit availability and latency SLOs. The production architecture layers multiple caches (serving 99.5% of static files and ~70% of all requests from cache) and the cutover used a proxy Worker that set a version cookie to route traffic between the legacy and new blog, with automatic fallback to the legacy site on 5xx errors, ramped 1% → 5% → 15% → 100% over launch day. The migration also shipped a frontend redesign (Kumo design system, dark mode) and two MCP servers (a Cloudflare Blog MCP and EmDash's own authoring MCP).

Key takeaways

  1. Customer Zero validated on real traffic and SLOs. Cloudflare ran its own blog on EmDash before external customers, treating the pain as a prioritisation signal. The scale target was concrete: normal ~75 RPS, spikes >5,000 RPS. (Source: this article; see customer-zero)
  2. k6 load-test scenario suite: ramp, breakpoint, burst. Three distinct scenarios — Ramp (gradually to 3× prod baseline, then cool down), Breakpoint (0→100 RPS over 10 min, stop when something breaks), Burst (immediate 7,000 RPS). The burst scenario used k6's constant-arrival-rate executor at rate: 7000, preAllocatedVUs: 4000, maxVUs: 10000.
  3. SLO thresholds encoded directly in the load test. Availability: fail if >0.01% of HTTP requests return 5xx. Latency: fail if p95 > 500 ms or p99 > 1000 ms. Expressed as k6 thresholds: http_req_failed: ["rate<0.01"], http_req_duration{status:200}: ["p(95)<500", "p(99)<1000"].
  4. Multi-layer caching carries the read load. Caches ordered by proximity to the user make the blog both fast and resilient — serving 99.5% of static files from cache and ~70% of all requests from cache, cutting database load and improving frontend performance.
  5. Proxy-Worker cookie-based cutover with fallback. A proxy Worker set a version cookie on requests, routing each request to the new or legacy experience, and fell back to the legacy blog on any 500 error. A NEW_BLOG service binding dispatched requests worker-to-worker (avoiding a public hostname → DNS → TLS → outbound HTTP round trip), reducing proxy latency.
  6. Progressive percentage rollout. Launch day ramped 1% → 5% → 15% → … → 100% of traffic, observing system health at each step and catching last-minute edge cases without impacting the majority of readers. Zero downtime was a non-negotiable requirement.
  7. Measurable performance gains. p95 response latency was flat and consistent under load on the new EmDash/Workers stack vs periodic spikes on the old platform; minimal errors while serving up to 850 RPS.
  8. Designing for agents. The migration shipped a Cloudflare Blog MCP server (search_posts, list_posts, get_post, list_tags) built in "a few hours" atop EmDash's APIs and AI search endpoints, plus EmDash's own authoring MCP for blog authors — no additional cost, part of the platform.
  9. Agents Week as first real test. 18 posts over 9 days, ~3M pageviews; the blog Worker served up to 450 RPS without issues and absorbed a 28,000 RPS DDoS attack (2026-08-10) via built-in DDoS protection.

Systems / concepts / patterns extracted

  • Systems: EmDash (CMS, Customer Zero subject), Astro (frontend rendering), Cloudflare Workers (runtime + proxy Worker + service binding), k6 (load testing), MCP (blog + EmDash authoring servers), Kumo design system (frontend redesign, not modelled).
  • Concepts: customer-zero (validated on real traffic), multi-layer-caching (99.5% static / 70% overall cache hit), concepts/core-web-vitals (Lighthouse-measured perf gains), concepts/cache-hit-rate, zero-downtime-migration.
  • Patterns: proxy-worker-cookie-based-traffic-routing (cookie-set proxy Worker routes legacy vs new, fallback on 5xx), load-test-scenario-suite-ramp-breakpoint-burst (three complementary k6 scenarios), patterns/staged-rollout (1%→5%→15%→100%), worker-to-worker-service-binding (NEW_BLOG binding avoids public-hostname round trip).

Operational numbers

  • Normal load ~75 RPS; spikes >5,000 RPS.
  • Burst load test: immediate 7,000 RPS for 1 min (4,000 preallocated VUs, 10,000 max).
  • SLO: <0.01% 5xx; p95 < 500 ms; p99 < 1000 ms.
  • Cache: 99.5% of static files served from cache; ~70% of all requests served from cache.
  • Post-migration: consistent p95 while serving up to 850 RPS.
  • Rollout: 1% → 5% → 15% → 100% over launch day (2026-08-12).
  • Agents Week: 18 posts / 9 days, ~3M pageviews, up to 450 RPS on the blog Worker; absorbed a 28,000 RPS DDoS on 2026-08-10.

Caveats

  • This is a migration/launch post with a strong product angle (EmDash promotion). The architecture content — load-test methodology, caching layering, proxy-Worker cutover, service bindings — is substantive and is what earns inclusion; the CMS-usability findings (scheduled-post bugs pre-v0.19.0, editor quirks) are product feedback rather than system design.
  • Cache-layer topology is described only qualitatively ("ordered top to bottom by proximity to the user"); the article's architecture diagrams are images not reproduced in the raw markdown, so exact layer identity (edge cache / tiered cache / Workers Cache / DB-adjacent) is inferred, not stated.

Source

Last updated · 766 distilled / 2,225 read