In the world of modern software development, the pull request is the cornerstone of collaboration. It is the crucible where code is scrutinized, refined, and ultimately merged into the collective body of a project. However, as software systems grow in complexity, so too do the changes required to maintain them. Broad refactors and critical migrations often demand massive, singular commits that cannot be cleanly broken down into smaller, "stacked" pull requests. This reality creates a daunting challenge for developers: navigating a pull request that spans thousands of files, contains over a million changed lines, and is buried under hundreds of inline review comments.
To ensure that the review process remains efficient and responsive under these extreme conditions, the team behind the GitHub Copilot app recently embarked on a complete reconstruction of its diff surface. The goal was simple yet technically formidable: create a user interface capable of rendering a million-line pull request with hundreds of active conversations without sacrificing fluidity or speed. By testing the new architecture against a massive open-source pull request—comprising 2,200 files and 400 inline comments—the engineering team has established a new standard for performance in large-scale code review.
The Scope of the Problem
Rendering a large diff at high speed is a well-understood problem in computer science. The standard solution is virtualization: the system renders only the rows visible on the user’s screen, plus a small buffer, while recycling DOM elements as the user scrolls. This creates the illusion that the entire document is loaded, maintaining a proper scrollbar size and scroll-to-row functionality, while keeping the actual memory footprint and DOM node count remarkably low.
However, the introduction of review comments fundamentally breaks the geometric assumptions that make code-only diffs fast. A code-only diff is predictable; every row is a line of code at a known font size, allowing the system to calculate the total height of the document before a single pixel is painted. Comments, by contrast, are notoriously unpredictable. Their height depends on markdown wrapping, the presence of reply boxes, expandable sections, and even whether images within the comment have finished loading. Because these factors can change dynamically, the traditional "all heights known before paint" contract fails.
Previous attempts to solve this involved using fixed-height slots for comments based on an estimate. Yet, in large pull requests, this approach consistently falls apart. An estimator that is accurate on average will inevitably create gaps of whitespace for some comments while forcing others to clip or sprout nested scrollbars. If the system attempts to measure the true height after rendering and writes that value back into the layout, it triggers a "scroll jump"—a jarring shift in the document position that occurs while the user is actively reading.
Two Geometries Instead of One
The breakthrough for the GitHub Copilot app team came when they realized they needed to stop forcing a single geometry to serve two distinct types of content. Instead, they split the document’s height into two independent domains: a deterministic "code geometry" and a dynamic "block geometry."
The code geometry remains the original, predictable world. It relies on exact, prefix-summed calculations that are never rebuilt, even when a comment thread is toggled open or closed. Conversely, the dynamic block geometry handles everything that cannot be predicted, such as review threads, draft comments, and reply composers. These blocks are anchored to specific file locations rather than fixed pixel coordinates, ensuring that even if a reflow occurs, the system never loses track of the content. By recording the width of these blocks during measurement and using fingerprints of their content, the system can determine the "effective height" of a block efficiently without forcing the entire document to recalculate.
This separation ensures that the performance of the diff surface is not bottlenecked by the most volatile elements on the page. Because the number of blocks is limited by the number of comments rather than the number of lines of code, the system remains performant even when the pull request is massive.
The Measurement Scheduler and Scroll Anchoring
The team’s initial approach to measuring these dynamic blocks was flawed: they experimented with using one ResizeObserver per comment block. While intuitive, this created a dangerous feedback loop where the observer would trigger a layout change, which in turn would trigger another observation, causing the cost to grow linearly with the number of mounted blocks.

The final solution implemented in the app is a single, idle- and scroll-gated measurement pass. This ensures that calculations are performed only when the system is not actively fighting the user for resources. Crucially, the team also had to solve the problem of scroll anchoring. When a measured height differs from an estimate, the system now corrects by identity rather than by pixel. It keeps track of what the user is currently viewing and adjusts the scroll position relative to that anchor. This prevents the "fighting the user" phenomenon, where a document suddenly shifts while the developer is in the middle of a sentence.
A significant hurdle involved distinguishing between user-initiated scrolls and internal scrolls caused by the system itself. For instance, toggling a file-tree sidebar might cause the diff to reflow. If the system incorrectly interprets this as a user scroll, it might disable necessary corrections, causing the content to drift off-screen. By strictly differentiating between user interaction and internal layout shifts, the team ensured that the interface remains stable and reliable.
The Pipeline Behind the Surface
The speed of the UI is ultimately dependent on the data pipeline feeding it. The Copilot app team adopted three key habits to optimize this flow. First, they prioritize the structure of the diff before the content, allowing the file tree and metadata to render while the document is still loading. Second, they defer per-item work—such as syntax highlighting or large markdown rendering—until those elements are near the viewport. This off-thread processing allows code to appear as plain text instantly, with colors and rich formatting applied as the resources become available.
Finally, the team refined their caching strategy. While they default to releasing large documents when a user navigates away to conserve memory, they implemented a selective cache that keeps the last few diffs resident. This provides an instantaneous experience when navigating between recently viewed files, while ensuring the app does not become bloated over a long session.
Mechanical Debugging and Autopilot
The complexity of this system meant that many bugs were invisible during normal testing. To address this, the team moved away from traditional console.log debugging, which relies on human observation. Instead, they built permanent, structured probes into the surface that assert the health of the rendering process in real-time.
These signals are integrated into an autonomous "change, measure, and improve" loop. The team utilized a headless probe that could drive the app through complex, declarative flows—such as scrolling to a specific fraction of a document or resizing the window—while automatically collecting data on rendering counts and performance bottlenecks. They also employed an "autopilot" system that drove the actual desktop app through these intense scenarios on a continuous loop, monitoring for gaps, clipped text, or layout errors.
By running this loop in their CI pipeline, the engineering team transformed performance from a manual, retrospective concern into an objective, automated metric.
Where This Leaves Us
The result of this engineering effort is a pull request view that handles massive, complex changes with the ease of a simple patch. A million-line diff now opens and scrolls with the responsiveness of a standard file view, and the conversation threads behave as a dynamic, fluid part of the document rather than a source of instability.
For developers who work on large-scale, high-velocity codebases, this represents a significant shift in how they interact with their tools. By treating the pull request not as a static document, but as a living conversation, the GitHub Copilot app team has demonstrated that the most effective way to handle scale is not just to optimize for speed, but to architect for the inherent unpredictability of the development process. For those who review code for a living, the difference is immediate: the tools no longer get in the way of the work.

