Skip to main content
  1. Posts/

The Recursive Turn: Andrej Karpathy Joins Anthropic's Pretraining Team

·7 mins
On May 19, 2026, Andrej Karpathy posted seven paragraphs to X announcing he was joining Anthropic. The post crossed 148,000 likes within a day, and someone on tech Twitter compared it to Kevin Durant joining the Warriors — the sport’s best individual player walking onto the roster that was already winning.

Karpathy’s own framing was quieter: “I think the next few years at the frontier of LLMs will be especially formative… excited to join the team here and get back to R&D.” He would lead a new pretraining-research group under Nick Joseph, Anthropic’s head of pretraining, built around a premise with a slight vertigo to it: use Claude to help build the next Claude.

The long way round #

The path here was not straight. Karpathy did his PhD under Fei-Fei Li at Stanford and taught the university’s CS231n, the course that trained a generation of computer-vision engineers on lecture notes as clear as any in the field. He was a founding research scientist at OpenAI in 2015, left in 2017 to run Tesla’s Autopilot vision team, returned to OpenAI in 2023 for a year of work on synthetic data and mid-training, and departed again in February 2024.

Five months later he founded Eureka Labs, an “AI-native school” pairing human-written courses with an AI teaching assistant. Its flagship product, LLM101n, is an open curriculum for building a ChatGPT-style model from scratch — in Python, then C, then CUDA — and it’s still live on GitHub, though the module list hasn’t visibly grown since 2024.

By mid-2026, reporters covering the Anthropic move noted he “hadn’t shared recent updates” on the startup since its launch. Eureka looks paused rather than closed — Karpathy has said he remains “deeply passionate about education” and intends to “resume my work on it in time.” For now, the school waits.

A decade of showing your work #

What makes Karpathy unusual in a field that increasingly runs on opaque, trillion-parameter systems is a parallel habit of building the smallest possible version of everything. The lineage runs microgradminGPTnanoGPTllm.cnanochat, each one a working model of the previous one’s ideas, stripped down until a single person could read the whole thing in an afternoon.

nanochat, released October 2025, is an ~8,000-line pipeline that trains a 561M-parameter chat model on 8×H100s for about $100 in compute. His pitch: “the best ChatGPT that $100 can buy.”

The same instinct drives his public writing. He coined “vibe coding” in a tweet on February 2, 2025 that racked up 4.5 million views — the idea of fully giving in to an AI coding agent’s suggestions and forgetting the code even exists — and then spent the rest of the year walking it toward something more serious: agentic engineering, the discipline of coordinating fallible AI agents while someone still keeps correctness, security, and maintainability intact.

Each additional nine of reliability costs about as much engineering effort as all the nines that came before it.

— Karpathy, on the “march of nines,” Dwarkesh Patel podcast, October 2025

He’d learned this the hard way already, at Tesla, trying to get Autopilot from “mostly doesn’t run a stop sign” to “never runs a stop sign.” On that same podcast he argued AGI is still roughly a decade away, drawing a line between the industry’s self-declared “Year of Agents” in 2025 and the actual decade of unglamorous reliability engineering he thinks it will take. Of some of the more breathless product claims circulating at the time, he was blunter still: “it’s slop.”

The company that ate 2026 #

Anthropic was founded in 2021 by siblings Dario and Daniela Amodei and five other OpenAI alumni, on the premise that safety research belonged at the center of frontier AI development rather than bolted on afterward. It trained Claude before ChatGPT existed publicly, and it has spent the years since compounding both its research reputation and, more recently, its balance sheet at a pace that looks less like a startup trajectory and more like a phase change.

DateMilestone
Mar 2025$61.5B valuation
Sep 2025Series F — $183B
Feb 2026Series G — $380B
May 2026Series H — $965B, Karpathy joins
Jun 2026Sonnet 5 and Fable 5 ship
Jul 2026Opus 5 ships

The model lineup jumped straight from the 4.x line to a new Claude 5 family across June and July 2026: Sonnet 5 as the cheaper, near-Opus agentic workhorse; Fable 5 as the most capable model Anthropic has released publicly, with some queries auto-routed to a more cautious model under the hood; Mythos 5, the same model gated behind an invitation-only research program; and Opus 5 arriving weeks later at a lower price than Fable 5 for nearly the same intelligence.

Meanwhile: a $1.5B copyright settlement over pirated training books, closed July 2026 — the largest of its kind on record — and a February directive from the Trump administration telling federal agencies to stop using Anthropic’s tools.

Underneath the model releases sits an infrastructure build-out sized to match: over $100B in committed AWS spend for Trainium capacity, commitments for up to a million Google TPUs, and by August 2026 a $10B compute deal with CoreWeave on top of it. Enterprise customers passed 300,000; revenue reportedly grew from roughly $9B annualized at the end of 2025 to somewhere between $30–47B by mid-2026. It is, by any measure, a company moving as fast as capital and hardware allow.

Why this hire, why now #

There’s a real tension sitting at the center of this move, and it’s worth naming rather than smoothing over. Karpathy has spent the past year as one of the industry’s most credible skeptics of AGI hype — the person willing to say, on the record, that the “Year of Agents” was mostly slop and that real reliability is a decade of grinding work, one nine at a time. He is now drawing a paycheck from the company raising money at a trillion-dollar clip on the opposite bet: that the frontier is close enough, and moving fast enough, to justify unprecedented capital and unprecedented urgency.

The way those two positions reconcile is in the specific shape of his new job. Anthropic isn’t hiring Karpathy to make bolder predictions; it’s hiring him to shorten the distance between “we have an idea for how pretraining should work” and “we have evidence for whether it does.” A pretraining team that uses Claude to generate hypotheses, design experiments, and build eval infrastructure is a bet that AI-assisted research — not just more GPUs — is what compounds fastest against OpenAI and Google. If each nine of reliability really does cost as much as all the nines before it, then the only way to afford the next one is to get more research done per unit of human attention. That’s the wager, and it’s one where his skepticism is arguably an asset rather than a contradiction: someone who doesn’t believe the hype is a reasonable person to put in charge of figuring out what’s actually true.

It also isn’t as large a swerve from his past decade as the “AGI skeptic joins the AGI company” framing suggests. Karpathy’s defining project, across Stanford, Tesla, and a string of from-scratch repositories, has been making neural networks legible — to students, to Twitter, to anyone patient enough to read a few thousand lines of code. Teaching a model to help design the experiments that build the next model is, in a strange way, the same project pointed at a new audience: instead of teaching a human how the machine works, he’s teaching the machine how to help build the next one.

He’s spent ten years explaining, line by line, how these systems work. His new job is to find out whether the system can start explaining itself, one nine at a time.