Back to feed

Claude made itself three times faster in two weeks: inside Anthropic's performance sprint

Anthropic made claude.ai and the desktop app three times faster in two weeks: over 3,000 changes, zero incidents, and deterministic measurement ratchets guarding every win.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — FsDUOUV9Vs8
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

In August, Anthropic made the core experience of the claude.ai website and the Claude desktop app roughly three times faster in a single two-week push. Users had complained about slowness for a long time, and the team openly agreed and got to work. According to the official engineering post on claude.dev, more than 3,000 changes shipped during the sprint without a single customer-facing incident. What surprised me most is that the entire effort ran from one Slack channel, with a model on duty in every thread.

Where to start

The team began by analyzing usage data and focused on four journeys covering 95 percent of all activity: launching the app, starting a conversation, loading an existing conversation, and sending a message. Across web, desktop, and products, those journeys became thirteen distinct measurements. At the 75th percentile, time to a typeable page on a fresh load fell from 3.1 seconds to 0.55, opening a new code session dropped from 0.8 to 0.3 seconds, and loading a cloud session went from 2.6 to 0.73 seconds. In aggregate, the team estimates tens of thousands of hours of waiting eliminated every day.

The sprint kicked off with about twenty hand-picked projects, and the model estimated each project's impact in milliseconds beforehand. By day three, twelve of thirteen targets were already met. To speed up launches, the composer was baked straight into the page HTML so users could type while React was still initializing. A precompiled V8 code cache was prepared for the desktop shell's main process, the composer stayed mounted between conversations, hovered sessions were prefetched, and wasteful sidebar re-renders were cut by 90 percent.

Anything measurable can be improved

The sprint's philosophy fits in one sentence: once the model can measure something, it can make it faster. Measurement used to be step zero: add a counter, wait for data, then understand the problem. Here measurement became the first step of the climb, and the team's highest-leverage work turned into finding new numbers to climb. An engineer named Sam proposed counting JavaScript instructions instead of wall-clock time, and within eleven minutes five separate measurement threads were running. This account is told in detail in the official engineering post on claude.dev, which shows the method is reproducible.

Wall-clock time is noisy and unusable as a CI gate, so the team built a ladder of deterministic counters : instruction counts via Valgrind for pure-JavaScript paths, plus React commit counts, V8 call counts, style recalculation counts, and DOM mutation counts for browser paths. Every new benchmark had two jobs: provide a number to move in the lab, and serve as a ceiling in CI that could only ratchet downward. Any benchmark that failed to prove it tracked wall-clock time was thrown out without mercy. According to the news summary on iphoneincanada.ca, this discipline was the core guarantee that let two weeks of work ship safely.

The proof experiment ran on two hot paths: the routine assembling a conversation's message tree and the scanner for status lines in code output. Profiling showed a quarter of the first path's instructions came from megamorphic dictionary lookups resolving the same message ID three times. An hour later, instruction counts were down 48 and 31 percent while wall-clock time had fallen 78 and 44 percent. Two new ratchets landed in CI with a daily job lowering each ceiling whenever the count dropped. As the analysis on analyticsindiamag.com points out, such deterministic locks are the most practical tool against gains rotting away in fast-moving codebases.

One hundred fifty parallel threads

The sprint quickly settled into a loop: someone opens a thread with a recording of a slow moment, the model traces the flow and builds a benchmark demonstrating it, returns with a few risk-sized pull requests once the lab looks promising, watches the deploy and reads field data, then tightens the ratchet on a win or flips the flag off and iterates on a miss. More than 150 threads ran at once, and roughly a third of all pull requests carried extra telemetry or guardrails with them. In the assessment published on fourweekmba.com, shipping over 3,000 changes without incident is presented as the strongest evidence for this loop discipline.

My favorite case was the sidebar jank hunt. Rows popped in at different times after page load and the page felt jittery, yet no existing monitor caught it. An engineer named Issac suggested looking directly at the browser's Layout Instability interface, and the model produced a telemetry event mapping each shift to a named region and phase. The test failed 20 out of 20 runs on the main branch and passed 20 out of 20 on the fix. Field data showed 31 percent of web loads shifted something after the page was usable with no user interaction, and culprits like a late header row, a caret sliding once the user name loaded, and a list jumping when the scrollbar appeared were cleaned up one by one.

Threads were left open instead of closed, and a single thread sometimes produced fifty or a hundred optimization requests. Increasingly the model, not the humans, opened new threads on its own. Shelley, an engineer on the team, called the model a numbers demon. It found 6,900 React hooks and 900 store subscriptions re-rendering on every keystroke in the typing path. It showed a single root selector adding 24 milliseconds to every DOM change, a forgotten command causing half a million hidden reloads a day, and identical cache snapshots cloned into IndexedDB twice a minute. According to the report on ciol.com, these findings rank among the most striking examples of measurement breeding new opportunity.

One of the most elegant fixes came from an em dash. Highlighting a finished code block froze the page for about a second, and the culprit was a single non-Latin-1 character in the reply. V8 stores strings containing such characters as UTF-16, which pushed every highlighting rule onto the slower two-byte path. The fix was a twenty-line change copying each code block into a one-byte string before highlighting. As the sprint recap on lavx.hu stresses, small but deep findings like this could never be caught without lab measurement.

No speed without guardrails

Because nearly every hot path was touched, safety was built up front: every request went through automated review with at least one human approval, tests were written before optimizations, and anything user-visible shipped behind a short-lived flag. Around two hundred flags were opened in two weeks and more than half were cleaned up before the sprint ended. The static composer was especially brittle; the real React component was rendered in a virtual document and drift was tested across fourteen viewport sizes within one pixel, while a keystroke test through the handoff forgave not a single lost character.

The most entertaining bug hunt began with a screen recording four hours after the internal release. The composer box dropped about 10 pixels when the page opened in a new tab. The model found the cause in Chrome's behavior rather than the team's code: on managed browsers the new-tab page carries a 56-pixel footer, typing in the address bar pre-renders the page at the shorter height in the background, and the page grows about 100 milliseconds after first paint, pushing the composer down. The model pinned the layout and added a test mimicking the pre-render flow. This case is told in the official engineering post on claude.dev as the finest example of closing the lab-field gap.

The humans' job: ambition, taste, direction

The loop was productive but not autonomous, and humans had three jobs. The first was ambition: the model tended to keep scope narrow, file tickets, and pad estimates, so the team pushed it to be bolder while trusting the guardrails. The second was taste: every user-perceptible change arrived with before-and-after recordings, and humans decided calls like whether a skeleton appears instantly or after half a second. The third was direction: each thread stayed deliberately narrow, and the conversation covered which surfaces to prioritize, how to merge colliding threads, and when to close one with diminishing returns. A 900-line request was rejected in one sentence as too complex for 2 milliseconds per send.

One side quest became legendary with its 8-millisecond rule. A frame-rate readout was added to a long streaming reply and the target was set at 120 Hz: 8.33 milliseconds per frame. Using the developer tools' frame control in a headless browser, the model achieved exactly 240 frame starts in 240 frames and walked the reply frame by frame. Length-dependent work vanished through memoized results of finished blocks, tokenizing logic for growing code fences moved to a worker, and tables revealed cell by cell. Nearly sixty requests landed in one thread, main-thread load for long replies fell from 750 to 200 milliseconds, and a 120 Hz laptop held full frame rate from start to finish.

Once the planned projects finished early, the team explicitly asked the model for wacky ideas and hill climbing entered a new round. The host validates this call from personal experience: a similar invitation once led him to fork a custom runtime and gain a two-to-fourfold speedup. As the analysis on fourweekmba.com notes, this second wave went far beyond the original targets and delivered the sprint's real leap. The lesson is plain: when targets are met, look for new things to measure instead of stopping.

Through the host's lens

The video's creator approaches the process cautiously as someone who has criticized Anthropic's engineering quality for years, and he parallels the post with his own product's performance campaign. Local-storage caching, an instantly appearing sidebar, and his model picker's loading lag keep echoing against the article. His core lesson: teach the model to measure first rather than asking it to fix performance. The sponsored segment in the middle, introducing a sign-in standard that agents can use on a user's behalf, is dispatched in a single sentence.

The sprint's aftershocks reached upstream too, with contributions landing in Electron, Chromium, and Node. Desktop and web are now about three times faster than in early August, with ratchets guarding the level. Yet the team says the job is unfinished: top-percentile latency, secondary flows, and extremely long threads still need work. As the report on iphoneincanada.ca also notes, users have started feeling the speedup in daily use. The channel is still open, and the team plans to keep going one thread at a time.

Visualization: nodesdaily AI

Key moments

  1. Opening: why claude.ai got three times faster
  2. Four key journeys and the p75 numbers
  3. Running the sprint from a Slack channel with Claude Tag
  4. Valgrind counters and the deterministic measurement ladder
  5. The 6,900-hook census and style recalculations
  6. Hunting the Chrome pre-render layout shift
  7. The 120 Hz target and the 8-millisecond frame budget
  8. Closing: measuring is the first step to improving

AI commentary

"What struck me most is how the team turned measurement discipline into product culture. My favorite detail is that slowness was treated as a number with an owner, never as an excuse. This approach looks copyable even for small teams."

AI assessment

The strongest counterpoint concerns the source of the numbers: the 3x figure, the 3,000 changes, and every percentile improvement come from Anthropic's own measurements, with no independent verification available. The analysis on fourweekmba.com makes exactly this caveat, noting that every claim is relayed from the company's engineering post dated September 23. The percentages look dramatic, yet the hardware and network conditions behind them are unspecified, so readers should not expect identical gains on their own machines.

There are gaps in scope as well: the team itself says the work is unfinished, with the 95th percentile and very long conversations still lagging. The mobile experience, low-bandwidth regions, and accessibility receive no mention in the post. The sprint's extraordinary two-week focus also leaves open whether the same discipline survives at normal product tempo. Ratchets may guard the wins, but nobody knows if they will be loosened under feature pressure.

The speaker's position deserves attention too: the host is a developer who has criticized Anthropic's engineering for years and sells his own coding product, so he reads the post to extract lessons rather than to praise it. That distance adds credibility to the retelling. The sponsored segment in the middle of the video, promoting a sign-in standard for agents, has nothing to do with the performance story, and compressing it into a passing mention is the right editorial call.

The practical takeaway for readers fits in three steps: first attach a number to the slowness, then make that number deterministic and lock it into CI, and deliberately keep the scope narrow. Starting with boring line items like hook counts and redundant paints instead of grand architectural dreams can produce a felt difference in two weeks. Teams adopting this method can use the loop described in the official engineering post on claude.dev as their template.

Sources

7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

claude · anthropic · performance · react · v8 · ai agents

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…