Rendering huge pull requests in the GitHub Copilot app¶
Summary¶
GitHub rebuilt the pull-request review surface in the GitHub Copilot app
so that an extreme diff — a real open-source PR with 2,200 files, >1 million
changed lines, and >400 inline review comments — opens, scrolls, and behaves
like a normally sized PR. The core insight: a code-only diff is fast because
every row is a line of code at a known height, so the whole scroll geometry can
be computed up front ("all heights known before paint"), but review comments
break that contract — a comment's height depends on markdown wrapping,
expandable <details>, an inline reply composer, and images that finish loading
after paint. The fix is to split the document into two independent geometries
(deterministic code height + estimated-then-measured dynamic-block height), a
single idle- and scroll-gated measurement pass instead of one
ResizeObserver-per-block feedback loop, and anchor-preserving scroll
correction that adjusts by identity rather than by pixel. Underneath, the data
pipeline streams structure before content, defers per-item work (syntax
highlighting off the main thread), and keeps a small resident cache of recent
diffs. Almost every bug was found not by eye but by an autonomous
change→measure→improve loop: permanent structured invariant probes, asserted
as CI budgets, driven by a headless probe lane and an unattended autopilot that
reads the app's own on-disk instrumentation.
Key takeaways¶
- Two geometries, not one. Total height =
deterministic code height (exact, prefix-summed, never rebuilt)+Σ dynamic-block effective heights (estimated → measured)+ scroll padding. Code rows keep the "all heights known before paint" contract; comments get a weaker, honest contract: heights are bounded, measured lazily, corrections are small and anchored to what the user is looking at (Source: sources/2026-09-23-github-rendering-huge-pull-requests-in-the-github-copilot-app). - Virtualization is table stakes for the code rows. Mount only the on-screen rows (~100 real DOM rows) plus a margin, recycle them on scroll, and drive scrollbar size / scroll-to-row from a typed-array height table. Because every code row is a known height, the table is computed once and never corrected.
- Dynamic blocks are keyed by identity, not pixel. Each review thread /
draft / reply composer is a block with a stable key, anchored to
file + line + side (not a pixel coordinate), carrying a fingerprint of
everything that changes its height (content,
<details>open state, composer active) and a width bucket so an ordinary window resize doesn't invalidate every measurement. Effective height = measured (if valid) → cached (if fingerprint+width match) → estimate. - One
ResizeObserver-per-block that writes height back into layout is the feedback loop to avoid. It can retrigger itself and its cost grows with every mounted block. GitHub designed this first and rejected it during performance hardening. - The measurement scheduler is a single idle- and scroll-gated pass. It runs
only when the visible range settles (never per scroll frame, never mid-scroll),
is scoped to ~2400px of the viewport (O(viewport), not O(document)),
reads all mounted candidates in one batch (single reflow, no interleaved
writes), does at most one off-screen render for a nearby unmounted block,
and skips blocks taller than the viewport (their over-reservation hides
below the fold). Per-block
ResizeObservers survive but by default only flag a block for re-read — they never write a height themselves. - One deliberate synchronous exception. For a resize the user caused
themselves (expanding
<details>, opening a reply, an image landing) on a mounted on-screen block, the observer measures and commits the correction in the same frame before paint, so the block grows and the code below it repositions together. Guardrails: at most one synchronous commit per frame and never during an active scroll. - Correct by anchor, not by pixel. Before applying height deltas, capture what the user is anchored to (a row or block, by identity) plus offset; apply deltas; resolve the same anchor to its new pixel position; scroll so it stays put. Blocks above the viewport adjust by the delta; content hydrating below the viewport does not (you can't see it).
- "Is the user scrolling?" must be a check your own side effects can't satisfy. A guard on "last observed scroll" was tripped by programmatic scrolls the surface itself emitted when the file-tree sidebar toggled and line-wrapped rows reflowed — so it skipped the very correction meant to keep your place. Fix: distinguish user scrolls from surface-caused scrolls.
- Pipeline: stream structure before content; defer per-item work. File tree
- metadata paint while the document still loads; the full set of review threads resolves up front (so no comment block is inserted after scrolling starts); syntax highlighting runs off the main thread (rows appear as plain text, colored later); large markdown / suggested-change context is built only as it approaches the viewport (Source: sources/2026-09-23-github-rendering-huge-pull-requests-in-the-github-copilot-app).
- Release diffs on navigate, but keep a small resident cache. Diff documents are large; holding every visited one leaks memory. But an instantly-drawn shell around an empty diff looks broken, so GitHub keeps the last few diffs resident, evicts beyond that, and lets a background refresh notice staleness.
- Find bugs mechanically, not by eye. The whitespace-under-comments class of bug is intermittent, engine-specific, and scroll-position-specific. GitHub built permanent structured probes answering invariants on every render (Is the surface viewport-bound? How many rows/blocks mounted? Is measurement coalescing to one commit/frame? How large are scroll corrections? Did any comment block get inserted after scroll started — must be zero? Do per-block observers tear down on unmount, or are we leaking one per block?) and asserted them as budgets in an end-to-end test against a synthetic many-comment huge-PR fixture.
- Put the loop on autopilot. A headless probe lane runs a declarative
JSON flow (open PR, scroll to a fraction, toggle a details block, resize the
window) against a mock server, reading React render counts, the performance
timeline, and a
requestAnimationFramejank sampler — so an agent can profile any flow described in plain English without editing source. An autopilot drives the real desktop app through the huge-PR flow unattended (cold then warm; toggling<details>, opening/cancelling composers, collapsing files, toggling the sidebar, sweeping deep into the file list, resizing), mirroring every measurement to the app's on-disk log with a per-sample health signal (a warm sample is healthy only if there are no unfilled gaps between comments, no blank comment blocks, and real thread content mounted across the whole scroll range).
Operational numbers¶
- Stress fixture: 2,200 files, >1,000,000 changed lines, >400 inline review comments (a real open-source PR).
- Virtualized DOM: ~100 real rows mounted at once regardless of total row count.
- Measurement scope: blocks within ~2400px of the viewport are measurement candidates.
- Scheduler budget: ≤1 synchronous height commit per frame; ≤1 off-screen render per nearby unmounted block; blocks taller than the viewport get 0 off-screen renders.
- Invariant budget asserted in CI: 0 comment blocks inserted after scrolling started once backend topology has landed; 0 leaked per-block observers after unmount.
Architecture notes¶
The code-geometry domain is the same design as GitHub's other diff surfaces: an imperative recycled code-row renderer (no React component per row), typed-array geometry for offset math, backend-owned diff documents streamed structure-first, and an imperative scroll API with exact "scroll to row N." The dynamic-block geometry is the new domain — a separate index so a resizing comment never forces the code geometry to rebuild, and the block count is bounded by comments (a few thousand) rather than by rows (a million). This is a sibling surface to the web Files-changed tab rewrite: both reach for list virtualization at the p95 tail and both hydrate/measure visible content first, but the Copilot-app surface adds the mixed-height comment problem that the code-only web-diff work did not have to solve at the same depth.
Caveats¶
- This is a client-rendering / frontend-architecture article — the "at scale" is DOM-node count, render frames, and scroll jank, not distributed systems. The reusable design vocabulary (virtualization, two-geometry split, anchor-preserving correction, off-the-hot-path measurement, streaming structure-first, invariant probes as budgets) is what generalizes.
- Numbers are single-fixture stress figures, not fleet-wide percentiles; the post does not disclose the diff-data store, the SSR substrate, or hardware.
- The autopilot / headless-probe loop is described qualitatively; no framework name or open-source release is given for the probe harness itself.
Source¶
- Original: https://github.blog/engineering/user-experience/rendering-huge-pull-requests-in-the-github-copilot-app/
- Raw markdown:
raw/github/2026-09-23-rendering-huge-pull-requests-in-the-github-copilot-app-078c955a.md
Related¶
- systems/github-copilot-pr-view — the surface this article describes.
- systems/github-pull-requests — sibling web Files-changed diff surface.
- concepts/list-virtualization — the row-recycling technique both surfaces rely on.
- concepts/streaming-ssr — stream-structure-before-content is the same discipline at the data layer.
- concepts/real-user-monitoring — the "measure real behavior, not hand-rolled logs" sibling.
- patterns/measurement-driven-micro-optimization — the change→measure→improve loop, here on autopilot.
- concepts/hot-path — measurement is kept off the per-frame hot path.