Judging whether a result is real
82Distinguishing a genuine effect from noise, a lucky seed, or a benchmark quietly leaking into training data requires exactly the skeptical judgment automated systems are worst at applying to their own output.
AI Researcher · Designs and tests the algorithms behind machine intelligence, in a field now racing to automate a growing share of its own research process.
Darker cells mean a higher score for this topic on that metric.
LessMore
Last reviewed Sources & creditsMedia creditsMethodology
Most days split between reading recent papers, writing and debugging training code, waiting on and analyzing results from long-running experiments, and discussing findings with collaborators. Contrary to the popular image, a large share of the job is diagnosing why a model is underperforming or a result won't replicate, not brainstorming new architectures from scratch.
For an independent research role at a university or a top industry lab, effectively yes — a PhD remains the field's real credential, since there is no licensing exam. Strong engineers without one increasingly enter through industry residency programs or a public record of open-source work and cited preprints, but titles like 'research scientist' still skew heavily toward doctorate holders.
Parts of it clearly are: literature summarizing, boilerplate experiment code and first-draft related-work sections are already routinely AI-assisted, and labs are openly experimenting with systems that propose and run their own experiments. Framing what a result actually means, and taking responsibility for a published claim, have proven much harder to hand over.
It varies enormously by employer and seniority. The US median for the broader 'computer and information research scientist' category was $140,910 in 2024, but new PhDs joining a large technology company's research lab often start well above that, and a small number of senior researchers at frontier AI labs have reportedly been offered total packages worth seven or eight figures during the 2024–2025 hiring competition.
The lines blur constantly, but a rough distinction holds: a researcher's output is usually a paper, a new method or an insight about why something works, judged by peer review; an engineer's output is usually a working system in production, judged by whether it ships and holds up under real use. Many people do both across a career.
Python dominates, paired with a deep-learning framework — PyTorch is now the most common in research, with JAX popular for large-scale and Google-affiliated work. Beyond code, researchers depend on shared GPU or TPU compute clusters, experiment-tracking tools like Weights & Biases, and arXiv, where most new results circulate as preprints months before formal peer review.
AI research is the profession most directly exposed to its own output. The tools it builds are now being turned, deliberately, on the research process itself — summarizing papers, proposing experiments, writing training code — which makes this section less hypothetical than in almost any other profile on this site.
That exposure does not make the job obsolete, but it does make the argument about what survives unusually concrete: a handful of labs are openly trying to build systems that run the entire research loop, and the honest answer, as of now, is that they can do parts of it but not the parts that matter most for judging whether a result is true.
A meaningful share of an AI researcher's routine work — literature triage, first-draft experiment code, hyperparameter search, drafting a related-work section — is already comparably fast for an AI system to do, and multi-agent research systems built explicitly to automate more of the loop are an active project at several major labs. What has not been automated is judging whether a result is real, deciding which question is worth asking, and being accountable when a published claim turns out to be wrong.
Scored from the tasks, not the job title. Lower is safer.
Jobs AI cannot take →Distinguishing a genuine effect from noise, a lucky seed, or a benchmark quietly leaking into training data requires exactly the skeptical judgment automated systems are worst at applying to their own output.
Deciding that a specific, hard, under-explored problem is worth months of a lab's compute and attention is a bet made with incomplete information — the part of research furthest from pattern-matching over past papers.
Someone has to stand behind a result when it fails to replicate or turns out to have a flawed baseline — a model cannot hold professional or scientific accountability for its own output.
Weighing a system's likely social effects, safety risks or dual-use potential before publishing draws on values and context an automated research pipeline has no grounds to reason about on its own.
The intuition for which surprising result is worth chasing, developed by watching hundreds of experiments succeed and fail over a career, is not yet something any system has demonstrated at a senior researcher's level.
Given tens of thousands of new preprints a year, tools that summarize and rank papers by relevance are already in routine use across the field.
Wiring together a standard training loop, data loader or evaluation script is now often faster to generate and check than to write from scratch.
Automated tuning systems can search a parameter space far more exhaustively and cheaply than a researcher manually trying configurations one at a time.
A first pass at summarizing prior work relevant to a new paper is increasingly AI-assisted, though the specific framing and honest comparison still need a researcher's check.
A growing share of a researcher's time goes to specifying what an automated system should try and judging its output, rather than writing and running every experiment personally.
As the field's own 'bitter lesson' predicts, access to large-scale compute increasingly determines which labs can even test a given idea, concentrating frontier research inside a small number of well-funded organizations.
As AI-assisted tools make it easier to produce a plausible-looking result quickly, rigorously verifying that result is becoming a more valued and more separately staffed specialty within research groups.
Systems designed to propose hypotheses, run experiments and draft papers with limited human oversight — projects like Sakana AI's 'AI Scientist' (2024) and Google DeepMind's 'AI co-scientist' (2025) — are themselves now an active area of publication.
Studies how to keep increasingly capable systems reliably doing what their designers intend, a specialty that barely existed as a distinct career path before the 2010s.
Builds and maintains the distributed training and compute infrastructure frontier research now depends on as much as any individual algorithmic idea.
Works at the boundary between technical research and government or corporate policy, translating what a lab's own systems can and cannot do into rules or safeguards.
Deliberately tests models and AI-generated research outputs for failure modes, security holes and false claims before they reach production or publication.
Three reversible lenses: augment the work, replace a slice, or open a niche. Teaching marks — not forecasts.
Keep the role; AI speeds drafts, triage, or research while judgement and accountability stay human.
A narrow task stack may compress first (templates, first drafts, routine scoring) while adjacent craft grows.
Oversight, integration, and domain QA roles can appear where AI output must be trusted in regulated settings.
Demand for AI research talent is very likely to stay strong through the next decade, but the entry-level version of the job is already changing: fewer roles will exist purely to run routine experiments by hand, and more will expect a new researcher to direct, question and verify work an AI-assisted pipeline produced almost immediately.
The clearest long-run advantage will sit with researchers who can frame a genuinely new question, spot when a plausible-looking result is quietly wrong, and take responsibility for a claim in public — the same judgment the field has always needed, now applied to a research process that increasingly writes its own first draft.
Closest neighbours on the six-score profile — not the same field only.
Designs and fabricates the transistors inside every computer, phone and weapon, using machines precise enough that only a few factories on Earth can run them.
AI-resistant 60 📦Builds the systems that train, deploy, monitor and govern machine-learning models in production.
AI-resistant 54 🩺Diagnoses illness and manages health for years afterward through examination and evidence, not a single operation — medicine's generalist and long-term guide.
AI-resistant 67 🧠Diagnoses and treats mental illness through medication and therapy, one of medicine's only specialties with legal authority to detain a patient in crisis.
AI-resistant 73 👔Advises clients, drafts the documents that bind them, and argues their case when it reaches court — carrying personal legal liability if the advice is wrong.
AI-resistant 58 🔐Protects systems, data and people by finding, preventing and responding to digital attacks.
AI-resistant 63Writes, tests and maintains the code that runs modern life — and is one of the first professions watching AI automate its own daily work.
AI-resistant 35 🛰️Designs, analyzes and certifies the aircraft, rockets and spacecraft that leave the ground, working to safety margins that leave no room for guessing.
AI-resistant 74 🌉The profession that turns rivers, rock and gravity into bridges, roads and clean water — civilization's quiet load-bearing trade since Imhotep.
AI-resistant 72 🔌Designs and fabricates the transistors inside every computer, phone and weapon, using machines precise enough that only a few factories on Earth can run them.
AI-resistant 60 🦾Designs the machines that sense, decide and act in the physical world, where the hard problem was never intelligence but the world itself.
AI-resistant 65 🔐Protects systems, data and people by finding, preventing and responding to digital attacks.
AI-resistant 63 📦Builds the systems that train, deploy, monitor and govern machine-learning models in production.
AI-resistant 54 🌬️Designs, builds and improves wind, solar, storage and grid systems that turn renewable resources into dependable electricity.
AI-resistant 72