Skip to content

SYSTEM Cited by 1 source

GitHub Copilot app — pull-request view

The GitHub Copilot app's pull-request view is a rebuilt code-review surface inside the GitHub Copilot app, engineered to keep the review experience fast when both the diff and its review conversation are enormous. It is a sibling to the web Files-changed tab but solves a harder rendering problem: a diff interleaved with hundreds of variable-height inline review comments. This page covers the [[sources/2026-09-23-github-rendering-huge-pull-requests-in-the-github-copilot-app|2026-09-23 engineering retrospective]].

Stress shape

The team opened the biggest PR they could find to pressure-test the design:

  • 2,200 files
  • >1,000,000 changed lines
  • >400 inline review comments

The goal: this extreme PR should open, scroll, and behave like a normally sized one — comments render in full (not clipped into a scrollable box), expanding a collapsed section moves the code below it and nothing else, and returning to a PR puts you where you were.

The core problem

Rendering a code-only diff at speed is well understood: virtualize the rows, keep the mounted DOM small (~100 real rows), and exploit the fact that every row is a line of code at a known height — so the entire scroll geometry can be computed up front and never corrected ("all heights known before paint").

Comments break that contract. A comment's height depends on how its markdown wraps, expandable <details>, an inline reply composer that grows as you type, suggested-change diffs, reactions, and images that change height when they finish loading — none of which are known until render time, and some of which keep changing after first paint. Reserving a fixed estimated slot per comment fails: it over-reserves most (whitespace gaps) and under-reserves the expensive ones (clipping / nested scrollbars), and measuring the real height afterward and writing it back shifts everything below while the user is scrolling — a scroll jump.

Architecture

Two geometries instead of one

total height = deterministic code height        (exact, prefix-summed, never rebuilt)
             + Σ dynamic-block effective height  (estimated, then measured)
             + scroll padding
  • Code geometry — the original deterministic world: imperative recycled code-row renderer (no React component per row), typed-array offset math, backend-owned diff documents streamed structure-first, imperative "scroll to row N" API. Never rebuilt when a comment resizes.
  • Dynamic-block geometry — everything unpredictable (review threads, drafts, reply composers). Each block has a stable key that survives content loading, is anchored to file + line + side (not a pixel), carries a fingerprint of everything that changes its height (content, <details> open, composer active), and records the width bucket it was last measured at. Effective height = measured (if valid) → cached (if fingerprint + width still match) → estimate. Blocks live in their own index; count is bounded by comments, not rows.

Measurement scheduler

The first design — one ResizeObserver per block that writes its measured height back into layout — was built and rejected: an observer that writes a height into the layout of the element it watches can retrigger itself, and cost grows with every mounted block. The shipped design is a single idle- and scroll-gated pass:

  • Off the hot path — runs when the visible range settles, never per scroll frame, and waits entirely while a scroll is in flight (a reflow mid-scroll is the jank being avoided).
  • Scoped to the viewport — only blocks within ~2400px of the viewport are candidates (O(viewport)); distant blocks ride their estimate until they approach.
  • On-screen reads win — reads every mounted candidate in one batch (single reflow, no interleaved writes); a mounted block is never skipped for a stale estimate. (This one rule fixed the worst bug: blank strips under comments, caused by a mounted block left on a too-tall estimate.)
  • Off-screen measurement is a bounded fallback — at most one off-screen render for a nearby unmounted block; blocks taller than the viewport skip even that.
  • Observers only flag, don't write — each mounted block keeps a ResizeObserver that by default just marks the block for the idle pass to re-read; it disconnects on unmount, and an inactive PR tab observes nothing.
  • One deliberate synchronous exception — for a user-caused resize (expanding <details>, opening a reply, an image landing) on a mounted on-screen block, the observer measures and commits the correction in the same frame before paint, so the block grows and the code below repositions together. Guardrails: ≤1 synchronous commit per frame; never during an active scroll.

Anchor-preserving scroll correction

When a measured height differs from its estimate the scrollbar arithmetic changes and the naive result is a viewport jump. The fix corrects by identity, not by pixel: (1) capture what the user is anchored to (row or block, by identity) plus offset; (2) apply height deltas; (3) resolve that same anchor to its new pixel position; (4) scroll so the anchor stays put. Blocks above the viewport adjust by the delta; content hydrating below does not; a <details>/reply you toggled in a visible block suppresses above-block correction so the interaction feels direct.

Sharp edge (bug): "don't correct while the user is scrolling" was implemented as a guard on the last observed scroll — but programmatic scrolls the surface emitted when the sidebar toggled (line-wrap reflow shifting the coordinate space) refreshed that timestamp, so the guard skipped the very correction meant to keep your place. Lesson: any "is the user interacting?" check must be one your own side effects can't satisfy.

Data pipeline

  • Stream structure before content — file tree + metadata paint while the document loads; the full set of review threads resolves up front (so no comment block is inserted after scrolling starts).
  • Defer per-item work — syntax highlighting runs off the main thread (rows appear as plain text, colored later); large markdown and suggested-change context build only as they approach the viewport.
  • Small resident diff cache — releasing a diff document on navigate is the right default (documents are large), but an instantly-drawn shell around an empty diff looks broken, so GitHub keeps the last few diffs resident, evicts beyond that, and lets a background refresh notice staleness.

Finding the bugs mechanically

Almost every bug was intermittent, engine-specific, and scroll-position-specific ("a strip of whitespace below some comments, but only sometimes, only on big PRs, and it heals if you scroll past and back"). So the surface carries permanent structured invariant probes answering questions on every render — Is the surface viewport-bound? How many rows/blocks mounted? Is measurement coalescing to one commit per frame, and how long does the frame take? How large are the scroll corrections? Did any comment block get inserted after scroll started (must be zero)? Do per-block observers tear down on unmount, or are we leaking one per block? — and asserts them as budgets in an end-to-end test against a synthetic many-comment huge-PR fixture.

The loop runs on autopilot:

  • Headless probe lane — a declarative JSON flow (open PR, scroll to a fraction, toggle a details block, resize the window) against a mock server, reading React render counts, the performance timeline, and a requestAnimationFrame jank sampler; an agent can profile any flow described in plain English without editing source.
  • Autopilot — drives the real desktop app through the huge-PR flow unattended (cold then warm; toggling <details>, opening/cancelling composers, collapsing files, toggling the sidebar, sweeping deep into the file list, resizing), mirroring every measurement to the app's on-disk log with a per-sample health signal (warm sample healthy only if no unfilled gaps between comments, no blank blocks, real thread content mounted across the whole scroll range).

Seen in

Last updated · 766 distilled / 2,225 read