It starts with a stubborn experiment: in the summer of 2014 a Stanford doctoral student sat down and hand-labelled some fifteen hundred hard photographs — lookalike dog breeds, half-hidden objects — to measure how far a human was still ahead of the machine, and on 2 September he published the verdict: 5.1% error for himself against 6.8% for GoogLeNet, a margin of one and seven tenths of a point. Everything else orbits that figure: Geoffrey Hinton's class in Toronto, the master's at UBC, the doctorate under Fei-Fei Li, the 22 October 2012 post saying we were «really, really far away», five years at Tesla fusing eight cameras into a single trunk, «vibe coding» in 2025, and pre-training at Anthropic from 19 May 2026 — the chronicle of a margin thinning out.

Fifteen hundred images, one at a time
In the summer of 2014 a Stanford doctoral student sits down in front of a screen and begins doing by hand what a machine does in seconds: look at a photograph and say what it contains. Not just any photograph — the hard ones, dog breeds that resemble one another, half-hidden objects. He labels roughly fifteen hundred of them 8. The reason is simple, and a little stubborn. Andrej Karpathy wanted to know by how much a human being was still beating the world's best image-recognition program, and no one had ever measured that margin properly. So he put himself on the test bench, working from the same list of a thousand categories the machine was given.
He published the result on September 2, 2014: human error 5.1%, GoogLeNet's error on the same sample 6.8% 8. One and seven-tenths of a point of advantage. It is the only moment in this story in which the protagonist appears as a competitor rather than an author — and, as we shall see, it is also a photograph taken at the last possible instant. Everything else turns on that number. Because Karpathy's career — from Geoffrey Hinton's class in Toronto to Anthropic's pre-training team, by way of five years running artificial intelligence at Tesla — is, seen up close, the chronicle of that margin thinning out. With a few widely circulated details that the documents do not confirm: we will say so where they come up.
Four numbers to get your bearings
gamma97
A résumé with its dates in order
Karpathy was born in Slovakia. He says so himself in an interview — “I was born in Slovakia and my family moved to Toronto when I was 15” 4 — and the Slovak press confirms it independently 6. The move to Toronto at fifteen is therefore documented in his own words. On the date of birth, however, something has to be said. October 23, 1986 circulates everywhere, nobody disputes it, and no primary source documents it: in the material we verified it traces back to encyclopedia entries and their derivatives, not to a registry or a statement by the man himself. This is not a suspicion: it is the difference between an attested fact and a repeated one.
At the University of Toronto he took a bachelor's degree with a double major in computer science and physics, plus a minor in mathematics — his personal website says so 1. That is where he encountered deep learningthe approach to artificial intelligence in which a computational network learns directly from examples, instead of having the rules programmed into it by hand: “This is where I first got into deep learning, attending Geoff Hinton's class and reading groups” 1.
The timing matters. He was attending Hinton's lectures and reading groups in the very years when Hinton's lab was about to change the field — we will see shortly how. That is not a merit of his; it is a matter of position: he was in the right room.
Then came the master's at the University of British Columbia, 2009-2011, advised by Michiel van de Panne, who works at UBC on reinforcement learning and physics-based simulation of movement 25. The subject was simulated robots: creatures that learn to move inside a world of equations before they move inside the world. And finally Stanford, 2011-2015, with Fei-Fei Li 2. The dissertation — “Connecting images and natural language” — is in the university library's catalog, with Percy Liang and Christopher D. Manning on the committee 3. The title already says where all of this was heading: teaching a machine not only to see, but to tell.
The career on a single time line, with the field's milestones beneath it
gamma97The leap that opened the decade
On September 30, 2012, while Karpathy is in his first year of doctoral work at Stanford, a neural network trained by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton wins ImageNet's annual image-recognition contest 7. It is called AlexNet. It does not win: it pulls away. The contest's yardstick is top-5 errorthe share of images in which the correct answer does not appear among the five labels the program flags as most likely — the lower, the better. AlexNet finishes at 15.3%. The runner-up is at 26.2% 7[rif:8bis]. That is 10.8 points of separation in a contest where competitors fought over a tenth of a point.
This is the number to keep in hand while reading the rest of this piece. Because two years later we will be talking about 5.1 against 6.8, and those two values look close to each other only if you forget that in 2012 the best in the world was wrong fifteen times out of a hundred.
The word going around back then was “dominated.” It is worth unpacking: it means exactly 10.8 percentage points on a task where annual improvements were counted in fractions. That is not emphasis, it is the measurement.
ILSVRC 2012, the gap in points
gamma97The measured core: how much was left to the human
This career has no revenue series to line up in a column. The only growth documented in numbers runs the other way: machine error on images falling, year after year. And inside that descent, one horizontal line — human performance — that nobody had ever traced precisely before the 2014 experiment.
Karpathy traced it and published it. On his sample of roughly 1,500 hard images, he was wrong 5.1% of the time, GoogLeNet 6.8% 8. Out of a hundred images, five errors of his against nearly seven of the machine's: a margin of 1.7 points, that is, fewer than two images in a hundred.
Here, though, the number must be handled with its own conditions attached, and it is the author himself who sets them. The 6.8% is GoogLeNet's error measured on his personal sample: on the full test set of a hundred thousand images that same network scored 6.7% 8. And the official result certified by the ILSVRC 2014 challenge, in the classification task, is 6.66-6.67% 11. The 5.1%, for its part, is not an official measurement: ImageNet does not certify human performance, and that value is the estimate of a single, heavily trained annotator on a sample he chose himself. Presented as “man beat machine in competition” it would be misleading — presented for what it is, a home-made measurement done with method and stated with its limits, it is more interesting still.
Because the real news is not who won. It is by how much: 1.7 points on a hard sample, when two years earlier the best in the world was missing 15.3 out of a hundred. The curve was falling so fast that the horizontal line was about to be crossed. Karpathy wrote it while he was measuring it. His prediction, in that post, was that before long human beings would be able to surpass the best classification models only with considerable effort, expertise and time 8. It was not a trophy: it was an expiration notice, signed by the man holding the stopwatch.
The machines' descent and the human line
gamma97“Really, really far away”
On October 22, 2012, three weeks after AlexNet's victory, Karpathy publishes on his blog a post whose title is already a position: The state of Computer Vision and AI: we are really, really far away 9.
The piece takes a photograph — a scene with some people, a balancing act, a visual gag — and shows how much you have to know about the world to understand it: who those people are, what happens if that weight shifts, why it is funny. No image classifier of the time could come close to anything of the sort. Here a clarification is owed to the reader, because history told after the fact tends to tidy things up more than the documents allow. That post does not mention AlexNet and cites no ILSVRC error rate from 2012: the only reference to ImageNet is generic 9. It is not the hot take on a victory three weeks earlier; it is an argument about the gap between classifying an image and understanding a scene.
The difference is not pedantry. A post written against AlexNet would be the story of a skeptic contradicted by the facts. A post written alongside AlexNet tells another story, truer and harder: that you can witness the biggest technical leap of the decade and still be right in saying that almost everything about understanding is still missing. Both things were right at once, and that is what makes the episode instructive. The machine had just learned to put the correct label on a photograph; understanding why that photograph was funny remained an open problem — and the post of the time said so with a candor that later reconstructions tend to smooth away.
Eight cameras, one single network
In June 2017 Karpathy joins Tesla as Director of AI and Autopilot Vision, reporting directly to Elon Musk 12. The hire was Musk's own doing: in a 2017 email later produced at the Musk v. Altman trial he wrote, “Andrej is arguably the #2 guy in the world in computer vision” 16. On that email, a clarification. The text is documented, the recipient is not: every available source identifies them only as “a Tesla vice president” 1617. Jim Keller was indeed Tesla's vice president for Autopilot in that period 19, but the identification remains an inference — no consultable document confirms it.
The technical problem Karpathy found on his desk is more interesting than any email. A car with autopilot has eight cameras arranged all around the body: three at the front with different apertures, four on the sides, one at the rear. Each sees a slice of the world, and the slices overlap at the edges. If each camera has its own separate processing, the edges bring the disaster you would expect: the same truck leaving one camera's field and entering the next one's can become two different objects, or vanish for an instant at the seam. Each branch decides on its own, and then someone downstream has to reconcile eight opinions.
The documented solution is the opposite: a single shared trunkthe first part of the neural network, the part that extracts visual features, used in common by all the inputs instead of being replicated for each one that receives the eight synchronized streams together and fuses them before deciding, with the various functions — recognizing vehicles, lane lines, traffic lights — branching off only at the end, like heads on a single body 18.
Here too, verification imposes a limit. The single-network architecture is described and matches what Karpathy has laid out in public talks, but the available source is a popular reconstruction, not a Tesla document: the material contains no attestation of the starting state — the separate processing — nor of the attribution of the shift to his leadership 18. The drawing below reconstructs the layout and the principle; the undocumented dimensions are declared as such.
The autopilot's eight cameras and the two architectures, at the same scale
gamma97What he does today
Karpathy leaves Tesla in July 2022, after five years: “It's been a great pleasure to help Tesla towards its goals over the last 5 years,” he writes 1520. The dates hold up — June 2017, July 2022 — while of a brief return to OpenAI after Tesla, often recounted, we found no trace in the verified material.
On July 16, 2024 he announces Eureka Labs, a company sitting between artificial intelligence and education 21. The site describes it as a school built natively around AI: the teacher designs the materials, an automated assistant scales them; first course, LLM101n 22.
On February 2, 2025 he coins two words that would enter the trade's vocabulary: “There's a new kind of coding I call ‘vibe coding’, where you fully give in to the vibes” 23. He is describing the programmer who stops writing every line and starts directing a system that writes them for him. This one is worth unpacking too. “Vibe coding” sounds light, but what it describes is a shift of role: from author of the code to client of the code. The man who coined it, for that matter, has said he feels further behind than ever as a programmer, despite his experience — a statement that remains his opinion, not a fact.
On May 19, 2026 he announces that he is joining Anthropic: “I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative” 24. A company spokesperson specifies that he will lead a team dedicated to using Claude to accelerate pre-trainingthe phase in which a language model learns from text before receiving any training on specific tasks — the most expensive and most decisive part research 2425.
Two things have been left out of this issue, and we will say which. The alleged $100 million signing bonus offered by Meta to OpenAI employees: the figure has a single origin, a statement by Sam Altman on a podcast of June 18, 2025 26, disputed by Meta CTO Andrew Bosworth — “the market's hot. It's not that hot” 27 — and by a former OpenAI researcher who called it “fake news” 28; Altman himself concedes that none of his people accepted it 29. Packages of a different nature remain documented, such as the $200 million reported for Ruoming Pang 30.
And the phrase “dramatically lower” with which AlexNet's result is sometimes dispatched: without the figures it is not information, with the figures — 15.3 against 26.2 — it becomes so. We have put in the figures, and left out the adjective.
The hierarchy of evidence: what holds this biography up, piece by piece
gamma97On Karpathy's technical worth the opinions on record are emphatic and come from people who had an interest in recruiting him. Musk, in the 2017 email produced in court: “Andrej is arguably the #2 guy in the world in computer vision after Ilya Sutskever” 16. It is the judgment of an employer justifying a hire poached from another organization, and it should be read for what it is: an interested assessment, not a ranking.
The source material also carries the judgment of an unidentified colleague, according to whom Karpathy is among the world's best computer-vision experts, perhaps the best of all. Since it can be attributed neither to a named person nor to a consultable document, we report it as an opinion without verifiable references.
Opinion reported as such, not verified by the editorial team.
What remains to be seen
One last image, and we stop here. In October 2012 a first-year doctoral student writes that we are “really, really far away” from getting a machine to understand a scene, and he writes it with the candor of someone who is not building a public position but thinking out loud.
Fourteen years later the same person works on the pre-training of a model that writes code on command, at a company that sells that model. Between the two scenes lie an experiment with fifteen hundred hand-labeled images, five years spent teaching a car to see with eight eyes instead of one, and two words that entered a trade's vocabulary. The question that remains is not whether he was wrong. He was right about what he wrote — the machines of the time classified images, they did not understand scenes — and being right lasted him less time than he imagined. It is a combination that often befalls those who work at the frontier: you see the limit clearly, you get the deadline wrong.
The real question is another one, and it concerns whoever is reading today: which sentence, written now with the same honesty and the same confidence in one's own instruments, is about to age in exactly the same way.
Curiosities from the notebook
Here are the four paragraphs:
At twenty-two he could solve a Rubik's Cube in under sixteen seconds. That is the very first image anyone has of him, long before any laboratory or announcement: a puzzle scrambled in someone else's hands, then rebuilt into six solid faces before a stopwatch could round up to seventeen. Nobody filmed it as a credential. It was simply a private kind of quickness, the sort of thing that never makes it into a biography unless someone happens to ask, and it turns out to be the first thing anyone thought to mention about him at all.
One day he was walking through a library, surrounded by shelves that did not seem to end. Rows folding into more rows, more books than one lifetime could hold, and somewhere in that walk a plain, almost embarrassing thought arrived: he wanted to read all of it, and he could not. Not wouldn't, could not, no matter how disciplined he became. The thought did not stay a complaint for long. It turned, quietly, into a different question: if he could not learn everything there was to know himself, maybe he could build something that could. He kept walking. The shelves did not end, but the question had already changed shape.
There is a photograph of a man standing on a bathroom scale, the sort every gym locker room has, and behind him, unnoticed, someone lifts a foot and quietly presses down on the platform to nudge the number up. That someone happens to be the President of the United States. The people standing around are already laughing, some of them visible again in a mirror behind the scene, in on the joke before the man on the scale is. To find it funny you would need to know a dozen things at once: who is joking, why the number matters, that a reflection is not a second crowd, and that no machine of that year could hold even one of them, let alone all seven.
At fifteen he left with his family, his sister included, for a country they had never lived in, giving up a life in Slovakia that had been, by his own account, comfortable. Not desperate, not a life anyone needed saving from: comfortable, which makes it a stranger thing to leave. He would later call it his family's leap of faith, and say that proving it was not wasted is what still drives his ambition, more than any single achievement could. Two people packed a comfortable life into suitcases so a son and daughter could have a different one, and neither of them, as far as the record shows, ever gets named.
Supporting the thesis
- The documentary solidity of the path is high and verifiable: the Toronto degree with Hinton's class comes from the personal website, the 2009-2011 UBC master's with van de Panne and the 2011-2015 Stanford doctorate with Fei-Fei Li from the university's institutional page, the dissertation from the library catalog. The technical context — AlexNet at 15.3% against 26.2% — comes from the original paper. On this skeleton there is no room for doubt.
Against the thesis
- Some much-quoted details hold up less well than the story wrapped around them. The man-versus-machine comparison is a personal experiment on ~1,500 images, not an official measurement: the certified ILSVRC 2014 figure is 6.66-6.67%, not 6.8%. The 2012 post mentions neither AlexNet nor that year's results, so the causal link often attributed to it finds no support in the text.
The verdicts
The Slovak origin is confirmed by Karpathy himself and by independent Slovak press. The date of October 23, 1986 and the city of Bratislava are not, however, confirmed by any primary source in the material gathered: they trace back to encyclopedia entries and their derivatives. Nobody disputes them, nobody documents them.
Karpathy states in the first person that he moved to Toronto at fifteen. Direct source, uncontested.
The personal website reports the bachelor's degree at the University of Toronto with a double major in computer science and physics and a minor in mathematics. First-hand primary source.
Karpathy writes that he got into deep learning in Toronto by attending Geoff Hinton's class and the associated reading groups. The correct name is Geoffrey, not Jeffrey as sometimes reported, but the substance holds.
AlexNet, authored by Krizhevsky, Sutskever and Hinton, won ILSVRC-2012 with a top-5 error of 15.3% against the runner-up's 26.2%: a 10.8-point gap.
Stanford's institutional page records the master's at the University of British Columbia in 2009-2011 with advisor Michiel van de Panne, a UBC faculty member in reinforcement learning and physics-based simulation of movement.
The Stanford doctorate (2011-2015) with advisor Fei-Fei Li appears on the institutional page and in the library catalog, which records the dissertation “Connecting images and natural language” with Percy Liang and Christopher D. Manning on the committee.
The post exists and is dated October 22, 2012, with the stated title. It should be noted that the text mentions neither AlexNet nor the ILSVRC 2012 results: the only reference to ImageNet is generic.
The post of September 2, 2014 describes the experiment: manual labeling of roughly 1,500 hard images and comparison with GoogLeNet, winner of ILSVRC 2014 in the classification+localization task. It is a personal exercise by the author, not an official contest certified by ImageNet.
The two figures are textually exact but not comparable with the official result: the 6.8% is GoogLeNet's error on Karpathy's personal sample, whereas on the full 100,000-image test set it is 6.7% and the certified ILSVRC 2014 challenge result is 6.66-6.67%. The 5.1% is an estimate of human error on ~1,500 images, not certified by ImageNet.
Karpathy is on record as a founding member of OpenAI, where he was a research scientist from 2015 to 2017; CNBC and TechCrunch in May 2026 describe him as an “OpenAI co-founder.” The primary page of the 2015 announcement proved inaccessible.
TechCrunch, on June 20, 2017, reports the hire as Director of AI and Autopilot Vision reporting directly to Elon Musk. The email produced at the Musk v. Altman trial confirms that the recruitment was Musk's own doing.
The text of the email is documented — “Andrej is arguably the #2 guy in the world in computer vision” — and reported by MIT Technology Review and other outlets from the May 2026 trial onward. The recipient, however, is identified by every source only as “a Tesla vice president”: none names Jim Keller. Keller was indeed VP of Autopilot in that period, but the identification remains an inference.
The single-network architecture that fuses the eight cameras on a shared trunk is documented and matches what Karpathy has laid out in public talks. What is missing is a source attesting the starting state — separate processing of the streams — and attributing the shift to his leadership: the available source is a popular reconstruction, not a Tesla document.
The five years at Tesla are confirmed: joining in June 2017, leaving in July 2022, with Karpathy himself speaking of “the last 5 years.” The return to OpenAI after Tesla is not, however, documented by any evidence in the material gathered.
Karpathy announced Eureka Labs on July 16, 2024 as an AI+Education company; the official site describes it as a school built natively around AI, with LLM101n as its first course.
The coinage dates to the post of February 2, 2025: “There's a new kind of coding I call ‘vibe coding’.” Independent semantic reconstruction traces the term back to that post.
The move to Anthropic's pre-training team was announced on May 19, 2026 by Karpathy himself and reported by TechCrunch and CNBC. An Anthropic spokesperson specifies that he will lead a team dedicated to using Claude to accelerate pre-training research.
The $100 million figure has a single origin: a statement by Sam Altman on a podcast of June 18, 2025, with no primary documents. Meta CTO Andrew Bosworth called it dishonest, former OpenAI researcher Lucas Beyer branded it fake news, and Altman himself concedes that none of his best employees accepted it. Multi-year packages of a different nature remain documented, such as the $200 million reported for Ruoming Pang.
References
- Andrej Karpathy — official personal website — https://karpathy.ai/
- Andrej Karpathy — Stanford Computer Science (institutional page) — https://cs.stanford.edu/people/karpathy/
- Connecting images and natural language (Stanford Libraries catalog) — — 2016 — https://searchworks.stanford.edu/view/11849345
- Training Deep Learning Models in a Browser: Andrej Karpathy Interview — https://www.datascienceweekly.org/data-scientist-interviews/training-deep-learni
- Michiel van de Panne | Computer Science at UBC — https://www.cs.ubc.ca/people/michiel-van-de-panne
- Šéf AI v Tesle, rodák zo Slovenska (independent Slovak press) — https://zive.aktuality.sk/clanok/147462/sef-ai-v-tesle-rodak-zo-slovenska-je-med
- ImageNet Classification with Deep Convolutional Neural Networks (Krizhevsky, Sutskever, Hinton) — — 2012 — https://proceedings.neurips.cc/paper/4824-imagenet-classification-with-deep-conv
- What I learned from competing against a ConvNet on ImageNet — — 2014-09-02 — http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-con
- The state of Computer Vision and AI: we are really, really far away — — 2012-10-22 — http://karpathy.github.io/2012/10/22/state-of-computer-vision/
- AlexNet — Wikipedia — — accessed 2026 — https://en.wikipedia.org/wiki/AlexNet
- Going Deeper with Convolutions (GoogLeNet paper) — — 2014 — https://arxiv.org/pdf/1409.4842
- ILSVRC2014 Results — — 2014 — https://image-net.org/challenges/LSVRC/2014/results
- Andrej Karpathy — Wikipedia — — accessed 2026-08-17 — https://en.wikipedia.org/wiki/Andrej_Karpathy
- Tesla hires deep learning expert Andrej Karpathy to lead Autopilot vision — — 2017-06-20 — https://techcrunch.com/2017/06/20/tesla-hires-deep-learning-expert-andrej-karpat
- Tesla loses top AI executive who led Autopilot vision team — — 2022-07-13 — https://techcrunch.com/2022/07/13/tesla-loses-top-ai-executive-who-led-autopilot
- Musk v. Altman week 1 — MIT Technology Review — — 2026-05-01 — https://www.technologyreview.com/2026/05/01/1136800/musk-v-altman-week-1-musk-sa
- Musk Testifies in OpenAI Trial — — 2026-05-02 — https://finance.biggo.com/news/202605021520_Musk-OpenAI-Trial-Testimony
- Tesla Autopilot Explained: HydraNet and Vision-Only Driving — — accessed 2026-08-17 — https://www.thinkautonomous.ai/blog/how-tesla-autopilot-works/
- Tesla's VP of Autopilot and chip guru Jim Keller is leaving — — 2018-04-25 — https://electrek.co/2018/04/25/tesla-autopilot-jim-keller-leaving-chip/
- Tesla's artificial intelligence director announces he's leaving — — 2022-07-14 — https://www.cnn.com/2022/07/14/business/tesla-karpathy-ai
- Andrej Karpathy on X — Eureka Labs announcement — — 2024-07-16 — https://x.com/karpathy/status/1813263734707790301
- Eureka Labs — official site — — accessed 2026-08-17 — https://eurekalabs.ai
- Andrej Karpathy on X — coining of “vibe coding” — — 2025-02-02 — https://x.com/karpathy/status/1886192184808149383
- OpenAI co-founder Andrej Karpathy joins Anthropic's pre-training team — — 2026-05-19 — https://techcrunch.com/2026/05/19/openai-co-founder-andrej-karpathy-joins-anthro
- Anthropic hires OpenAI co-founder Andrej Karpathy, former Tesla AI leader — — 2026-05-19 — https://www.cnbc.com/2026/05/19/anthropic-hires-openai-cofounder-andrej-karpathy
- Sam Altman says Meta offered OpenAI staff $100 million bonuses — — 2025-06-18 — https://www.cnbc.com/2025/06/18/sam-altman-says-meta-tried-to-poach-openai-staff
- Meta's CTO Says Sam Altman Is 'Being Dishonest' About $100 Million Signing Bonuses — — 2025 — https://www.entrepreneur.com/business-news/meta-cto-sam-altman-dishonest-for-100
- Former OpenAI researcher Lucas Beyer pours cold water on $100 million Meta signing bonus — — 2025-06-26 — https://finance.yahoo.com/news/former-openai-researcher-lucas-beyer-001932926.ht
- OpenAI's Sam Altman says Meta offered employees $100 million sign-on bonuses but none have accepted — — 2025-06 — https://www.barchart.com/story/news/32946200/openais-sam-altman-says-meta-offere
- Meta's Hiring Spree Raised Compensation for Top AI Engineers and Executives — — 2025 — https://www.deeplearning.ai/the-batch/metas-hiring-spree-raised-compensation-for
- Anthropic hires OpenAI co-founder Andrej Karpathy, former Tesla AI leader (CNBC) — — 2026-05-19 — https://www.cnbc.com/2026/05/19/anthropic-hires-openai-cofounder-andrej-karpathy