A Princeton professor on AI and work at ICML: how to adapt as an individual
- Princeton CS professor Arvind Narayanan's ICML 2026 keynote in Seoul, "What will be left for us to work on?", builds on his and Sayash Kapoor's AI as Normal Technology framework.
- Across ~24 months of SAGE observations, frontier models from three labs leapt in capability; the composite reliability metric (consistency / robustness / calibration / operational safety) rose only 5–10 percentage points.
- In software's decide–execute–deliver stack, AI mainly compresses the middle execute layer (once ~1/3 of hours); the ends stay—or grow.
- ATM, radiology, translation, and software tools history: automation rarely cuts headcount one-for-one. Software employment rose ~10,000× across successive ~10× tool leaps.
- He splits RSI, human-level AI, economically transformative AI, and superintelligence into four non-entailing dimensions—rejecting "lab milestone = humans immediately out of work."
- Personal adaptation: push the ceiling, don't coast on the floor; balance productivity / growth / control; refuse black boxes; master first, then amplify; reinvest ~10 hours a week of saved time into skills.
Two competing stories: even AI people are fretting about jobs
Last week at ICML 2026 in Seoul, Princeton computer science professor Arvind Narayanan delivered a keynote titled "What will be left for us to work on?"
He splits the path into two—not a philosophy debate, but a practical stance each person has to choose:
His warning: if you bet on displacement and the world is amplification instead, you may miss history's best window to level up. The world is watching how AI people respond—if practitioners roll over and hand work to AI, the political backlash could get uglier than today's.
AI adoption runs through four stages—and the slowest one has barely begun
"Normal" in AI as Normal Technology does not mean AI is a hammer or a toothbrush. They treat it as industrial-revolution-scale tech. The point is a causal frame: how capability becomes economic and social impact, step by step.
Past work on electricity-style tech often uses diffusion of innovation: invention → innovation (downstream products like appliances) → diffusion (gradual adoption). They stretch that into four stages, using software engineering as the running example:
Stage four is the slow one. Even in software engineering—relative early adopters of coding agents—he says real org-level redesign has barely started. He allows a speculative beat: if agents can reliably ship multi-million-line codebases with few bugs and security holes, building one uniform product for a billion users gets less compelling. Software may tilt toward extreme per-team customization, and even "do we still need software companies as an org form?" becomes fair game. Those shifts are human and organizational, and history usually measures them in decades.
Electricity in factories: drop-in replacement never worked
Before electricity, factories ran on one giant steam engine, powering the plant through gears and belts. When electricity arrived, owners first swapped a generator for the steam boiler—hoping for a more efficient drop-in replacement. That path failed.
People often say agents will be drop-in replacements for human workers. Electricity's lesson: the real payoff came from reorganizing work, not swapping humans for machines one-for-one. And that was not the utility company's job—same for AI. Org redesign is not something AI labs finish alone. In their four-stage frame, it is the slowest stage, and today it has barely begun.
Measured: capability rose; reliability did not keep pace
Across industries, a huge gap sits between what AI could do and what people actually deploy. Slow adoption is part of it. They also suspect deployers hit hard walls that leaderboards never measure—before the rest of the industry does.
Reliability is the concern people name most often. SAGE collapses roughly a dozen reliability metrics into four dimensions:
General-purpose + high-stakes + fully autonomous still looks like pick two of three. Collaborative agents will keep outperforming fully autonomous replacements. Scaffolding and post-training should differ for each—not force both under one "more autonomy is always better" story.
Software engineering: writing code was never the bottleneck
A 2019 paper already argued that writing code is not software engineering's bottleneck. The past year, blogs rediscovered that. Coding agents sped up the middle layer—the whole job did not shrink by the same factor.
Machines do the cognitive heavy lifting; humans still run the job. The role becomes "operate this machine," not "hand-finish every unit of cognitive labor." Forklifts and cranes did not kill construction sites—they rewrote who does what.
The lump-of-labor fallacy: automation rarely cuts jobs one-for-one
The lump-of-labor fallacy treats work as a fixed pie: if AI takes a slice, jobs vanish forever. History often says otherwise—when efficiency rises, demand and job structure shift with it.
Software itself: from machine code onward, successive ~10× productivity-tool leaps came with ~10,000× more jobs—because total code to write grew faster still. He is not saying "no one will ever lose work." He is saying "automation rate = unemployment rate" is a bad extrapolation.
If recursive self-improvement actually arrives: four dimensions people mash together
Many labs openly chase recursive self-improvement (RSI). He takes that path seriously—without equating a lab milestone with "humans immediately have nothing left to do."
Early explorers could call a whole archipelago "Hawaii." Once the ship is close, failing to name the islands muddles where to sail next. RSI, human-level AI, economically transformative AI, and superintelligence often get chained as automatic dominos. He wants them separated.
Example: curing cancer is often bottlenecked by trials with thousands of subjects over 10–15 years—external constraints more FLOPs in the lab will not erase. Mapping compute or model milestones straight onto "society immediately out of work" misses those walls.
Why AI creativity lags—and how open-world evaluation measures it
Deep learning already excels at perceptual representations. Representations that support creativity and higher-order reasoning, he argues, still trail humans. A few hypotheses from cognitive science and practice:
SAGE's open-world evaluation: give an agent a few thousand dollars of budget plus a real ML problem a human expert already spent months on and wrote up—but not yet on arXiv—then have those same experts grade the output. The team has already run evaluations like "have an agent independently build and ship an iOS app," and is recruiting senior researchers to expand.
How should individuals adapt in this wave?
Framework and evidence first. In the second half—"Personal reflections on adapting to AI"—Narayanan does not hand out one answer. He opens his research workflow: how he is riding this capability surge, for the audience to check against their own.
First, pick a stance: displacement vs amplification
Those opening stories become two life configurations at the personal level:
Floor vs ceiling: where saved time goes
He frames personal strategy with a pair of words:
His move: when AI clearly lifts productivity, reinvest the saved time in long-term growth—new skills and workflows that complement AI. He spends roughly 10 hours a week only on learning and trying new processes.
"If I finish a day completely unexhausted, I did it wrong—I offloaded too much to AI and traded long-term growth for short-term output."
Three-legged stool: productivity · growth · control
Faster delivery with AI is not the only goal. Three legs have to stand at once:
Push only productivity and starve growth, and a rising floor flattens you. Chase only growth and ignore delivery, and reality ends you. Have both but surrender control, and long-term you become a button-pusher.
Two heuristics for keeping control
Machines do the cognitive heavy work; people stay in the cab. The job becomes operating, understanding, and controlling the machine—not hauling every brick by hand. Personal adaptation means not climbing out of the cab to become just another brick on the schedule.
Closing vision: human–machine co-superintelligence
In one sense, economically transformative AI has already started: not when some AGI switch flips, but as slow variables—reliability, integration, tacit knowledge, regulation—rewrite work. He rejects geopolitics reduced to "whoever hits a capability milestone first takes all the economic returns."
On superintelligence he stresses: tasks often have ceilings; human intelligence leans hard on learning and tools, and AI is another tool—so the contest looks more like AI-augmented people vs AI acting alone. If we default to a future where AI owns companies and decides hiring and firing, then lean only on alignment as a backstop, that safety posture "opposes safety more than it supports it."
He likens computers to "bicycles for the mind" and AI to "cranes for the mind"—lifting human potential to heights once unimaginable. The learning curve is steep, like a treadmill that never stops. He still treats co-superintelligence as a fight worth fighting: not abandoning work, but redefining it as a higher-ceiling dance with AI.
The capability floor rises on its own; you have to push the ceiling. Adaptation is not dumping all work on AI—it is reinvesting saved time in complementary skills, and staying in the control seat. Based on Arvind Narayanan · ICML 2026 keynote
A Princeton professor cools the hype: AI capability surges, reliability lags—and work is redefined, not erased
SAGE tracked models from three frontier labs over ~24 months: capability jumped; reliability rose only 5–10 percentage points
↓ One page · one animated chart
Even AI people are fretting about jobs
Arvind Narayanan, Princeton CS professor, keynoted ICML 2026 in Seoul last week. He runs SAGE Lab's AI agent evaluation research and co-authored the AI as Normal Technology framework. The talk faces the career anxiety of AI researchers and programmers themselves.
In 24 months, capability jumped; reliability barely moved
SAGE folds a dozen-plus reliability metrics into four buckets—consistency, robustness, calibration, operational safety—then scores models released by three frontier labs over ~24 months on two complementary benchmarks.
Programmers aren't vanishing—writing code was never the bottleneck
A 2019 paper already said writing code is not software engineering's bottleneck. Narayanan splits knowledge work into three layers: decide (set direction), execute (code and debug), deliver (own the release).
Productivity gains rarely cut jobs one-for-one
Economists call it the lump-of-labor fallacy: treat total work as fixed, so AI taking a slice permanently deletes jobs. History often runs the other way.
Banks opened more branches and hired more tellers for what machines can't do
Radiology employment grew; doctors are adopting AI
Translator headcount roughly flat; expected stable another decade
Software engineering employment rose ~10,000×
RSI does not automatically snowball into superintelligence
Many labs say they are chasing recursive self-improvement (RSI: AI helping improve the next AI). Narayanan's "Hawaii problem": once the ship nears the archipelago, you need island names—or you get lost on where to sail. He splits RSI, human-level AI, economically transformative AI, and superintelligence into four non-entailing dimensions.
- RSI — the spectrum is huge: upgraded AutoML on one end, replacing the creative sum of the world's AI researchers on the other. Which end companies mean is often unclear.
- Human-level AI — strong on verifiable tasks (speed, efficiency); still far behind on unverifiable creative work.
- Economically transformative AI — bottlenecks are reliability, system integration, domain tacit knowledge, and regulation; decades of gradual adoption chew them down.
- Superintelligence — many tasks have ceilings (weather chaos limits; 10–15-year drug-trial protocols). Humans already sit near those limits.
pouring out
am I done?!
doing everything now?
Reliability +5–10pp
only
or any task fails ~30%?
Most tests never split these
not job loss
- × ATMs "kill" tellers—banks hire more, open branches
- × Radiology "gone in 5 yrs" (Hinton)—employment rose
- × Translators "replaced"—near-human MT a decade, headcount steady
- × Devs "tooled out"—~10× tools, ~10,000× more people
I still have work now
~24-month window · 3 labs only
