Anti-debate // ICML 2026, Seoul
Notes from an anti-debate on whether we should keep funding a human research workforce once AI is better at research than we are.
Last week I did an anti-debate with Ludwig Schmidt at an AI-for-research session at ICML in Seoul. It was not recorded, at the participants' request, so what follows is a reconstruction of my side: the argument I made, the evidence behind it, and where I landed by the end.
The motion
"Even after AI surpasses humans in research ability, humans should continue to fund a significant proportion of today's human research workforce to conduct research."
I argued the proposition. The format was not designed to produce a winner. The anti-debate structure rewards honesty and updating: each side opens with a personal story, gives openings and rebuttals, and then, halfway through, argues the opposing case, states what would change its mind, and works toward synthesis. The effect is to turn the exercise into a search for where the disagreement actually sits.
We fixed the terms first. "Surpasses humans in research ability" was defined operationally: a double-blind board of expert reviewers, after real effort, prefers the model's research questions and results, or at least matches the best a human can produce. "Significant proportion" was set at roughly 80 percent for me and capped at about 20 percent for the opposition. The scope was ML research, roughly what ICML covers, with no wet lab. On those terms the disagreement was never about zero versus some. Both sides granted that the number is not zero and that the composition of research work changes substantially. What remained was a small set of parameters: how substitutable human research is for machine research, whether verification can be decoupled from doing, how large the un-automated fraction of the pipeline is, how correlated the failure modes are, how long a cut takes to reverse, and who ends up owning the returns.
My opening was meant to be personal. I came into AI as a skeptic, and I have spent much of my career red-teaming these systems and finding where they break. However, more recently, I have been shocked by how fast these systems are getting better at research, and I feel confused about what that means for the future of my job. That is part of why I said yes to arguing this side.
Three questions keep surfacing for me. How do you reconcile the two extremes: the Bay Area conviction that everything is about to change and the calmer view I hear elsewhere? What do you tell students who are in the messy middle and have to make career decisions this year? And how do we keep our own judgment sharp while the tools we lean on get sharper?
The first question is a large part of why I have opinions here. I spend enough time inside the San Francisco discourse to be tired of the doom that has settled over it, and of the frontier-lab framing that there is one correct path and that anyone on a different one is simply behind. It surfaces most often in response to claims that scientific progress is about to be mostly a function of which lab points its compute at your field, and that talent will drain out of academia accordingly. A great deal of exciting work is happening, and I would not argue otherwise. But we would not be where we are without academia, and there is something worth defending in the pursuit of curiosity for its own sake. That discourse is out of scope for the motion. It is the context in which I built the case.
The case rests on five parameters, in roughly the order they matter, from the economics of substitution through to insurance and the supply chain.
Substitutability: is machine research a substitute for human research, or a complement?
The debate turns on a single parameter, the elasticity of substitution between human-made and machine-made research. If the two are close substitutes, then once the machine is cheaper and better, spending moves to it and the human share collapses. If they are complements, cheap machine research raises the value of the human input. Research is the second case, for three reasons.
Consider first the distinction between a task and a job. The motion evaluates a bounded artifact, the questions and results a panel can score. Funding buys a bundle: choosing which problem to attack, running and interpreting an experiment, coordinating a group, communicating the result, and standing behind it. Automating one component of that bundle does not dissolve the rest. The assumption that it does is the lump-of-labour fallacy, the premise that there is a fixed quantity of research so that any part the machine takes is subtracted from humans. The historical record runs the other way. For most of the last century the share of national income paid to labour held within a narrow band, roughly 60 to 65 percent, one of Kaldor's stylised facts, which Keynes called one of the best-established regularities in economics, across a century of automation panics. The share has drifted down a few points since the 1980s, which is real and worth watching, but no wave of mechanisation produced the collapse the substitution story predicts. When a task inside a job is automated, demand for the complementary human tasks around it has usually risen.
Second, returns to the time horizon. Here, one piece of direct evidence comes from this community. In RE-Bench, METR compares sixty-one expert ML researchers with frontier agents on seven open-ended research-engineering problems. Given a two-hour budget, the best agent scores about four times the human experts; given thirty-two hours, the humans move ahead and score roughly double the best agent. The agents are fast, and on short, well-scoped problems they win. What they do not yet supply is the human return to long-horizon, open-ended effort, which is the regime research funding exists to support.
Third, the serial structure of the pipeline, which is Amdahl's law in economic form. Chad Jones formalises it as a weak-links, or Baumol, argument: when tasks are complements, total output is governed by the tasks that remain un-automated, and their relative price rises as everything around them gets cheap. Automating an entire category of work, even perfectly, raises aggregate output by a bounded amount, because the un-automated remainder becomes the binding constraint. The remainder is the cluster of research tasks with no verifiable target: setting the agenda, judging whether a result is real, communicating it, taking responsibility for it, and aligning it to human values, the trust and oversight layer of the whole enterprise. Those are the weak links, and their price rises as execution gets cheap.
Underneath all three sits the finding by Bloom, Jones, Van Reenen and Webb that ideas are getting harder to find. Sustaining Moore's law now takes roughly eighteen times the researchers it required in the early 1970s, with research productivity per person falling around seven percent a year. A world in which discovery is getting harder wants more capable searchers at the frontier, human or machine. Substitution requires research to be a fixed pie handed to a cheaper producer. It behaves like an expanding frontier that rewards adding capacity.
The mechanism that ties this together runs opposite to the substitution intuition. Jevons' paradox: when something becomes cheaper to use, total demand for it tends to rise rather than fall. Cheaper research does not shrink the appetite for research; it expands the number of questions worth asking and experiments worth running, each of which still needs a person to choose it, judge it, and stand behind it. Arvind Narayanan's ICML keynote makes this case directly, tracing it through the same automation episodes, ATMs and radiologists, where employment grew rather than collapsed once the task was automated. The human-held share of a project falls while the number of projects rises, and there is no reason to expect the product to shrink toward a skeleton crew.
The verification–doing coupling: you cannot audit research you have stopped practicing.
Grant the machine superhuman output. Someone still has to decide whether to believe a result, sign for it, and answer for it when it is wrong, and that capacity does not stand on its own. The supply of people who can verify research is downstream of the supply of people who do it: a verifier who can catch a subtle error in a superhuman system is someone who has done the underlying work long enough to know what a real result looks like. The oversight layer cannot be a small cadre of watchers detached from practice. It has to be drawn from, and kept inside, a working research pipeline.
Judgment is a practiced skill, and it decays without use. In a 2025 Lancet study, nineteen experienced endoscopists lost detection skill after three months of AI assistance: when the tool was removed, their unaided adenoma detection rate fell from 28.4 percent to 22.4 percent. Three months eroded expert perception in specialists with thousands of procedures behind them. A research workforce that stops running its own experiments for a decade does not only shrink the pool of referees; it degrades the capacity of the ones who remain to referee at all.
Demand for oversight rises from both ends at once. Automated researchers generate far more experiments, code, and decisions per unit time than any fixed set of humans can review, which raises the volume to be checked, while cutting the workforce shrinks the pool qualified to check it. The review layer of our own field already shows the strain. A 2024 study by Liang and colleagues estimates that between 6.5 and 17 percent of the text in peer reviews at recent AI conferences was substantially produced by language models. That is a demand curve for practiced human referees.
Alignment sits inside this layer, and it is the part I find hardest to imagine automating. Deciding whether to trust a high-stakes claim that has no clean ground truth is a judgment about values and consequences, and it cannot be delegated to the system under evaluation without circularity, because the optimiser cannot certify itself. My own research is on privacy and contextual integrity, where the ground truth is a contested, outcome-based judgment about which norms apply and whose interests count, and where no benchmark settles the question. Those judgments are made well only when the people affected are represented among the people making them. Aligning a system to human values is not a step you run once and store; it is a standing negotiation that requires humans to do the negotiating, and it does not have a fixed answer to hand to a model.
The pipeline is itself the oversight infrastructure, and we already treat training this way in other domains. In 1914 a Sperry autopilot flew hands-off, and by 1947 a US Air Force aircraft crossed the Atlantic, takeoff to landing, entirely under automatic control; large aircraft still carry two trained pilots today, because regulation and hard experience treat the human as the accountable agent when the automation reaches its edge (history of flight automation). We still teach children arithmetic in a world full of calculators for the same reason: the training exists to build the judgment that lets a person catch a wrong answer, not to compete with the tool on speed. Teaching research, and funding the people who do it, is how a field keeps producing verifiers who can still recognise when the machine is wrong; remove the doers and within a generation there is no one left qualified to referee. The deskilling result above is the near-term edge of exactly this problem.
Correlation of failure, and the value of a plural workforce.
Aggregating many judgments improves on a single one only when the errors are independent, a condition that goes back to Condorcet's jury theorem. Copies of one model do not meet it. Across more than 350 models, Kim and colleagues find that when two of them answer a question wrong, they choose the same wrong answer about 60 percent of the time, and the correlation is highest among the largest and most accurate systems, even across different providers and architectures. Adding a tenth model to a panel of nine does not add a tenth independent judgment; it restates a shared prior. Ensembling buys little when the members converge underneath.
Human researchers fail in less correlated ways, because their formation is not shared: different advisors, languages, subfields, and failures produce different blind spots. There is evidence that this heterogeneity is not just present but productive. Analysing 1.2 million US doctoral theses, Hofstra and colleagues find that researchers from underrepresented backgrounds introduce novel conceptual combinations at higher rates; across nine million papers, AlShebli and colleagues find that ethnically diverse teams produce higher-impact work. The diversity of a large research community works as error-correction: it covers regions of the hypothesis space that a correlated system leaves dark. The matching result on the machine side is that an algorithmic monoculture can lower the quality of collective decisions even when the shared model is individually more accurate (Kleinberg and Raghavan, 2021). Concentrating the world's research judgment into a few correlated systems is that monoculture at the largest possible scale.
There is a further concern about concentration, and it is about power rather than accounting. If research is fully automated, the capacity to set the agenda, to decide which questions get asked and which results are trusted, moves from a distributed public system of universities and agencies to the small number of firms that own the frontier models and the compute they run on. That capacity is currently spread across many institutions and countries, which is part of why science can correct itself and serve interests beyond any single owner. A publicly funded research base keeps the questions plural and the answers accountable to more than whoever controls the datacenter. This is the sense in which the motion is a normative claim and not only an empirical one, and it is why the concentration point is central rather than decorative.
Taste and serendipity: the parts of research that resist optimisation.
Grant that execution is solved. Two things stay scarce on the supply side.
The first is taste: choosing what is worth working on before the field agrees that it was. This resists the current training paradigm, because the reward signals available, citations, acceptance, downstream use, are all lagging measures of consensus. Optimising against them produces work aimed at what the field already values, which is the opposite of the contrarian bet that defines an important direction. The recent reasoning gains are consistent with this reading: reinforcement learning on verifiable rewards mainly sharpens a model's ability to reach answers its base model could already sometimes reach, and whether it expands the frontier of new directions is contested. For now, the taste for the unproven direction is supplied from outside the reward loop.
The second is exploration. A large, uncoordinated population of researchers with idiosyncratic paths is a wide portfolio of low-probability bets, most of which return nothing and a few of which are penicillin. Optimisation prunes the low-probability branches, because in expectation they look like waste. Some of that population works on questions with no near-term payoff, in astrophysics or pure mathematics, because they want to understand something, and that motive keeps the search broad in a way a returns-maximising allocation would not. The point of funding people to follow curiosity is that you cannot know in advance which unpromising line becomes the important one.
There is also a demand-side scarcity, and it is measurable. In controlled experiments (Horton and colleagues, 2023), people assign a lower value, including a lower willingness to pay, to an identical artifact once it is labelled machine-made, even when they cannot tell the difference. Around research this is beginning to appear as a premium on human-authored work. It is evidence that human-made research is a distinct good with its own demand, which is the substitutability claim from the first argument arriving from the consumer side. The uncertainty-of-outcome hypothesis from sports economics, that people pay to watch a contest precisely because a human might fail, has survived seven decades of testing. A helicopter reached the summit of Everest two decades ago, and the number of people paying to climb it, and the price of the permit, have only risen since. The result was never the whole of the value.
Reversibility, insurance, and the supply chain.
The final argument holds even if the first four fail. Part of the case for a large human research base is insurance, and insurance is priced by two quantities: how correlated the primary system's failures are, and how long the backup takes to rebuild once it has been cut. Both point the same way, and both are worth spelling out.
Start with correlation, because the AI research stack is correlated at every layer. The models are trained on substantially the same corpus, so a systematic gap or bias in that corpus is a systematic gap everywhere at once, with no independently-trained population to catch it. Training models recursively on model-generated output degrades them across generations, a failure mode documented as model collapse, and it is arrested only by a continuing supply of fresh human data. The physical layer is a single point of failure of a different kind. The large majority of leading-edge chips are fabricated on one island, and much of the world's oil still moves through one strait, so a disruption in the Taiwan Strait or the Strait of Hormuz would not be a local shock to be routed around; it would propagate through a network that everyone now depends on simultaneously. The pattern is local robustness with global fragility, the same structure that makes a genetic monoculture efficient right up until one pathogen finds the shared weakness.
Now reversibility, where the asymmetry bites. A human research base cannot be switched off and on. The path from a first-year PhD student to an independent investigator runs the better part of a decade, and the tacit knowledge that makes a working scientist, the parts learned by apprenticeship and never written down, disperses far faster than it can be regrown. A version of this is happening now: when a large share of a national research budget is frozen or cut in a single year, thousands of grants lapse and a substantial fraction of the affected scientists begin planning to leave, and capacity that took a generation to build is damaged in a single budget cycle. You cannot repurchase it at short notice at any price, which is exactly why an uncorrelated, slow-to-rebuild backup has to be held before it is needed rather than reconstituted afterward.
The premium is unusually cheap, and it returns a profit even if the insured event never occurs, because public research has a large measured social return on ordinary grounds. Fieldhouse and Mertens estimate the social return to non-defense public R&D at roughly 140 to 210 percent, and attribute about a fifth of postwar US business-sector productivity growth to it, concluding that it is substantially underfunded. The amounts are also small relative to what they are weighed against. The entire annual NSF budget is on the order of nine billion dollars, which pales beside the roughly half a trillion dollars of AI capital expenditure now spent in a single year. The choice is not human salaries against compute; compute already outspends the whole public-science enterprise by an order of magnitude. Maintaining the human base is, in budget terms, a small hedge with a positive expected return.
Taken together, that is what supports a figure near 80 percent, rather than any single argument. I was explicit in the room about what the number means. It refers to the proportion of people, reshaped into new roles, not to the share of the budget, which I expect to tilt heavily toward compute. A workforce can be most of the headcount and a small, shrinking fraction of the spending at the same time. Eighty is shorthand for most of them, doing different work, and it is defensible down to wherever the arguments above stop holding.
I will keep this compact so as not to strawman it, but there was a real case on the other side, and its economic core is the argument that moved me most. Every verifiable cognitive task we have learned to cheapen has eventually been automated, and once it was, we stopped paying people to do that specific task. ML research is unusually easy to check against ground truth, more so than almost any other science, so it is a strong candidate to automate early rather than late. Public money carries an opportunity cost measured in the discoveries it could have funded elsewhere, so parking most of it in human headcount once the machine is better is not obviously stewardship; it can be inertia. Keep a lean core of accountable people, retrain the rest with real support, and let the funding follow the returns.
Where the opposition landed, and where we converged, was that research should be reshaped rather than preserved as it is: much more AI-for-science, far more leverage per person, and fewer people executing the parts machines now do well. The remaining disagreement was how much of today's ML field survives on its own terms, versus being absorbed into a smaller, science-facing enterprise.
The agreement was larger than the disagreement. We converged first on reshaping. Neither of us wants today's field frozen in place; both of us expect roles to change and move around, with fewer people doing what machines now do well and much of the field shifting toward AI-for-science and adjacent disciplines, where the human contribution is choosing the direction, judging the result, and designing what to ask. The work migrates more than it disappears, and a great deal of it migrates toward the sciences that AI is now accelerating. We agreed too that more people should be doing science rather than fewer, and that the case is not only instrumental. Some of the most valuable research is done because someone wants to understand something, in fields with no near-term product, and a society rich enough to automate its labour can afford to fund that curiosity as an end in itself.
We also agreed, and I think this is underweighted, that governance is a shared priority rather than an afterthought. As research automates, who oversees it, who is accountable for its outputs, and how its benefits are distributed become research problems in their own right, and both of us think that work is under-supplied today. Two threads of it came up directly. One is dissemination: a result is not useful until it is understood and trusted, and that trust is built by scientists explaining their own work rather than handed to intermediaries who translate it at second hand. The other is liability: institutions still run on a named human who signs and is accountable, the engineer who seals the drawing, the investigator of record on a trial, the author who answers for the paper, and a model cannot occupy that role, because accountability is imposed on a person. Oversight, dissemination, and accountability are the parts of research that grow in importance as execution gets cheap, and they are where a reshaped workforce concentrates.
One thing did change for me over the hour. I walked in assuming that the work I most wanted to protect had to stay inside ICML and the other ML venues. I am no longer sure it does. Much of what I have been defending here, the science itself and the work of explaining it and building trust in it, may not belong in a machine-learning conference at all. It might be better served by its own home, a separate field, a separate conference, or a funded venue built for the purpose, rather than carried by the ML community as a side function. The case for funding the human work does not rest on that work wearing the ICML badge, and by the end I thought the two questions, whether to fund it and where it should live, are worth holding apart.
Thanks to Fanny for organising and moderating, and to Ludwig Schmidt for a really insightful discussion. Thanks to the friends and colleagues who argued these ideas with me beforehand and humoured me while I worked them out: Jan Leike, Charlie George, Alex Imas, Brian Jabarian, Ashish Vaswani, Arthur Kosmala, Eric Zelikman, and Yuchen He. And thanks to Sergei Gukov and Tim Gehrunger for the discussion afterward.