Judging whether a result is real
82Distinguishing a genuine effect from noise, a lucky seed, or a benchmark quietly leaking into training data requires exactly the skeptical judgment automated systems are worst at applying to their own output.
AI Researcher · Designs and tests the algorithms behind machine intelligence, in a field now racing to automate a growing share of its own research process.
AI research is the profession most directly exposed to its own output. The tools it builds are now being turned, deliberately, on the research process itself — summarizing papers, proposing experiments, writing training code — which makes this section less hypothetical than in almost any other profile on this site.
That exposure does not make the job obsolete, but it does make the argument about what survives unusually concrete: a handful of labs are openly trying to build systems that run the entire research loop, and the honest answer, as of now, is that they can do parts of it but not the parts that matter most for judging whether a result is true.
A meaningful share of an AI researcher's routine work — literature triage, first-draft experiment code, hyperparameter search, drafting a related-work section — is already comparably fast for an AI system to do, and multi-agent research systems built explicitly to automate more of the loop are an active project at several major labs. What has not been automated is judging whether a result is real, deciding which question is worth asking, and being accountable when a published claim turns out to be wrong.
Scored from the tasks, not the job title. Lower is safer.
Jobs AI cannot take →Distinguishing a genuine effect from noise, a lucky seed, or a benchmark quietly leaking into training data requires exactly the skeptical judgment automated systems are worst at applying to their own output.
Deciding that a specific, hard, under-explored problem is worth months of a lab's compute and attention is a bet made with incomplete information — the part of research furthest from pattern-matching over past papers.
Someone has to stand behind a result when it fails to replicate or turns out to have a flawed baseline — a model cannot hold professional or scientific accountability for its own output.
Weighing a system's likely social effects, safety risks or dual-use potential before publishing draws on values and context an automated research pipeline has no grounds to reason about on its own.
The intuition for which surprising result is worth chasing, developed by watching hundreds of experiments succeed and fail over a career, is not yet something any system has demonstrated at a senior researcher's level.
Given tens of thousands of new preprints a year, tools that summarize and rank papers by relevance are already in routine use across the field.
Wiring together a standard training loop, data loader or evaluation script is now often faster to generate and check than to write from scratch.
Automated tuning systems can search a parameter space far more exhaustively and cheaply than a researcher manually trying configurations one at a time.
A first pass at summarizing prior work relevant to a new paper is increasingly AI-assisted, though the specific framing and honest comparison still need a researcher's check.
A growing share of a researcher's time goes to specifying what an automated system should try and judging its output, rather than writing and running every experiment personally.
As the field's own 'bitter lesson' predicts, access to large-scale compute increasingly determines which labs can even test a given idea, concentrating frontier research inside a small number of well-funded organizations.
As AI-assisted tools make it easier to produce a plausible-looking result quickly, rigorously verifying that result is becoming a more valued and more separately staffed specialty within research groups.
Systems designed to propose hypotheses, run experiments and draft papers with limited human oversight — projects like Sakana AI's 'AI Scientist' (2024) and Google DeepMind's 'AI co-scientist' (2025) — are themselves now an active area of publication.
Studies how to keep increasingly capable systems reliably doing what their designers intend, a specialty that barely existed as a distinct career path before the 2010s.
Builds and maintains the distributed training and compute infrastructure frontier research now depends on as much as any individual algorithmic idea.
Works at the boundary between technical research and government or corporate policy, translating what a lab's own systems can and cannot do into rules or safeguards.
Deliberately tests models and AI-generated research outputs for failure modes, security holes and false claims before they reach production or publication.
Demand for AI research talent is very likely to stay strong through the next decade, but the entry-level version of the job is already changing: fewer roles will exist purely to run routine experiments by hand, and more will expect a new researcher to direct, question and verify work an AI-assisted pipeline produced almost immediately.
The clearest long-run advantage will sit with researchers who can frame a genuinely new question, spot when a plausible-looking result is quietly wrong, and take responsibility for a claim in public — the same judgment the field has always needed, now applied to a research process that increasingly writes its own first draft.
Writes, tests and maintains the code that runs modern life — and is one of the first professions watching AI automate its own daily work.
AI-resistant 35 🛰️Designs, analyzes and certifies the aircraft, rockets and spacecraft that leave the ground, working to safety margins that leave no room for guessing.
AI-resistant 74 🔌Designs and fabricates the transistors inside every computer, phone and weapon, using machines precise enough that only a few factories on Earth can run them.
AI-resistant 60