Skip to content

CONCEPT Cited by 7 sources

Memory safety

Definition

Memory safety is the property that a program cannot access memory it isn't authorized to — no use-after-free, no buffer overrun, no double-free, no uninitialized-read. In C / C++ this is the programmer's responsibility; in Rust (post-borrow-checker), Java/Kotlin, Go, etc. it is enforced by the language/runtime.

The practical importance for distributed-systems engineering is that the vast majority of new exploitable vulnerabilities in C/C++ systems stem from memory-safety bugs (use-after-free, bounds errors).

The key data point: new bugs come from new code

The Android team published a 2024 analysis (link) showing that the majority of newly discovered memory-safety vulnerabilities come from recently-written code, not from the aged core codebase. The operational consequence:

If you want to prevent memory-safety bugs, you don't need to rewrite the world — you just need to stop introducing new memory-unsafe code.

This reframes the trade-off from "huge C/C++ rewrite" to "adopt a safe language at the margin, for new work."

Case study: Aurora DSQL's Postgres extensions

The systems/aurora-dsql team initially assumed its Postgres extensions should be in C — lowest impedance mismatch with Postgres's C API surface. In a code review they found several memory-safety bugs in a seemingly simple data-structure implementation. Rethinking against the Android data:

  • Postgres's C core is battle-tested over decades — not where new bugs come from.
  • The team's new extension code would be the source of almost all their new memory-safety bugs.
  • Rust for extensions → no loss of control / performance, + elimination of that new-code bug class.

Even though Rust extensions interact tightly with Postgres's C API, the team built Rust abstractions that encode invariants into the type system that Postgres relies on conventions / comments for. Classic example: a C char* paired with a len field, whose safe-use contract lives in header comments. In Rust, a single String type encapsulates the invariant — making "violating the invariant" uncompilable rather than "commented-against."

Why Rust (vs. other safe languages) for this class of problem

  • Keeps GC off the hot path → compatible with concepts/tail-latency-at-scale constraints.
  • Zero-cost abstractions → high-level code compiles to efficient machine instructions; no runtime penalty for using safe abstractions.
  • C FFI ergonomics → can call and be called by existing C codebases (Postgres, libc) with careful wrappers.

Managed-language alternatives (Java, Kotlin, Go) deliver memory safety but not the GC-free property needed by DSQL's data plane.

Trade-offs

  • Learning curve: "with Rust you have the hangover first." DSQL's JVM-native team lost a few weeks early on, then shipped 10× the TPS of their tuned Kotlin version.
  • Unsafe blocks: necessary at C-FFI boundaries. The win is that they are localized and reviewable rather than pervasive.
  • Ecosystem: at the DSQL pivot to Rust control plane, the team's Rust internal libraries had matured enough (AWS Auth Runtime's Rust client was faster than its Java counterpart) to make the move viable — an earlier attempt would have been blocked by missing libraries.

Seen in

  • sources/2025-05-27-allthingsdistributed-aurora-dsql-rust-journey — rationale for writing DSQL's Postgres extensions (and eventually its whole codebase) in Rust rather than C; Android team's "new code is where new bugs live" data invoked as supporting evidence.
  • sources/2025-07-17-datadog-go-124-memory-regression — the other end of the memory-safety trade-off: Go delivers memory safety but via a managed runtime, and a 2025 mallocgc refactor silently committed ~20% more RSS on real fleets without any program-visible symptom. Memory-safe ≠ memory-observable-from-inside-the-program; see go-runtime-memory-model.
  • sources/2024-05-31-dropbox-testing-sync-at-dropbox-2020 — Dropbox Nucleus chose Rust for the sync-engine rewrite; the post promises a follow-up specifically on how Rust's type system encodes invalid-state-unrepresentable invariants (e.g. "a node cannot exist without a parent"). Same Android-data-motivated family as Aurora DSQL's Rust extensions: memory-safe language for new code, even tightly interoperating with existing systems.
  • sources/2025-01-11-google-nearly-all-binary-searches-and-mergesorts-are-broken-2006 — sibling bug-substrate: integer-overflow in C/C++ (signed-overflow undefined per ISO/IEC 9899 §6.5) is one of the ways memory-safety bugs manifest — malloc(n * sizeof(T)) overflowing size_t produces an undersized allocation that the code then writes past. Bloch's 2006 binary-search post is the canonical discipline-level retrospective on the same class of hazards.
  • sources/2026-01-28-meta-rust-at-scale-an-added-layer-of-security-for-whatsapp — client-side cross-platform library instance. Meta's WhatsApp security team rewrites wamedia — the cross-platform C++ media library — from 160,000 LoC C++ (excluding tests) to 90,000 LoC Rust (including tests), with performance + runtime-memory advantages, and names this "the largest ever deployment of Rust code to a diverse set of end-user platforms and products that we are aware of." Ships monthly to billions of devices across WhatsApp, Messenger, and Instagram on Android/iOS/Mac/Web/Wearables. Meta's explicit 3-pillar strategy: (1) minimise attack surface, (2) invest in C/C++ assurance for what remains, (3) "default the choice of memory safe languages, and not C and C++, for new code." Motivation for wamedia specifically: "media checks run automatically on download and process untrusted inputs" — canonical memory-safe-language-for-untrusted-input instance. Rewrite methodology: parallel implementation with differential-fuzzing + integration + unit tests — see parallel-rewrite-with-differential-testing. Disclosed costs: initial binary-size tax from the Rust stdlib, build-system investment for diverse platforms. Extends memory-safety framing from server-side (DSQL) + managed-runtime (Datadog Go) + sync-engine (Nucleus) axes to the client-side cross-platform library axis.
  • sources/2026-05-18-cloudflare-project-glasswing-what-mythos-showed-us — AI-vulnerability-research triage-cost axis. Cloudflare's run of Mythos Preview across "more than fifty" of their repos surfaces a new memory-safety tax that compounds beyond the well-known exploit-surface tax: "We saw consistently more false positives from projects written in memory-unsafe languages." Memory-unsafe codebases produce more plausible-looking bug surfaces (bounds checks, pointer arithmetic, lifetimes) for AI vuln scanners to flag with hedged findings, inflating triage cost (see signal-to-noise-in-ai-vulnerability-triage). Combined with exploit chain construction capability — which reclassifies low-severity memory bugs as chain-eligible primitives — memory-unsafe substrates now carry (1) direct exploit-surface tax + (2) AI-vuln-triage tax + (3) chainability-amplification tax. Strengthens the wiki's memory-safety thread with a chain-aware-AI-triage data point.
  • sources/2026-09-01-databricks-collaboration-makes-us-all-stronger — shipped-OSS-extension axis. A memory-safety bug in the address_standardizer sub-extension of PostGIS — a caller-controlled grammar-rule value indexing a fixed-size internal array without a bounds check → out-of-bounds access — reachable by a normal tenant role on managed Postgres (Lakebase, Neon). Same "new/edge code in a battle-tested C core is where the bug lives" pattern as Aurora DSQL's extensions, now on the ubiquitous third-party extension axis: because PostGIS ships almost everywhere, "a memory-corruption vulnerability in a widely deployed extension is effectively a memory-corruption vulnerability in PostgreSQL itself." The lesson here is less about language choice than about ownership + containment: the platform owned the exposure, a microVM boundary contained the blast radius (no cross-customer impact), a downstream patch lever protected customers ahead of upstream, and the fix was driven upstream. Also a reminder that a real memory-safety fix can land upstream with no CVE, as a "minor" cleanup.
  • sources/2026-09-18-cloudflare-saving-another-100tb-of-ram-with-math-and-rust — struct-layout / safe-packing axis (not an exploit-class instance). Cloudflare shrank a hash Point in systems/pingora-ketama from 8→6 bytes but hit Rust's alignment rule: a struct's size is rounded to a multiple of its largest field, so changing the index field u32→u16 alone saved nothing. The tempting #[repr(packed)] is controversial for memory-safety reasons (creating references to unaligned fields is UB); the team instead stored the fields as a raw [u8; 6] array with typed getters (u32::from_ne_bytes / u16::from_ne_bytes), which compiles to the same code without the unaligned-reference hazard. Illustrates that memory-safe languages still expose low-level layout control — and that the safe idiom can be equivalent in codegen to the unsafe shortcut. See concepts/consistent-hashing for the broader optimization.
Last updated · 766 distilled / 2,225 read