Skip to content Skip to a section

🤖AI & The Future

AI Researcher · Designs and tests the algorithms behind machine intelligence, in a field now racing to automate a growing share of its own research process.

At a glance
Score intensity

Darker cells mean a higher score for this topic on that metric.

Last reviewed Sources & creditsMedia creditsMethodology

Quick answers

What does an AI researcher actually do day to day?

Most days split between reading recent papers, writing and debugging training code, waiting on and analyzing results from long-running experiments, and discussing findings with collaborators. Contrary to the popular image, a large share of the job is diagnosing why a model is underperforming or a result won't replicate, not brainstorming new architectures from scratch.

Do I need a PhD to become an AI researcher?

For an independent research role at a university or a top industry lab, effectively yes — a PhD remains the field's real credential, since there is no licensing exam. Strong engineers without one increasingly enter through industry residency programs or a public record of open-source work and cited preprints, but titles like 'research scientist' still skew heavily toward doctorate holders.

Is AI research itself at risk from AI?

Parts of it clearly are: literature summarizing, boilerplate experiment code and first-draft related-work sections are already routinely AI-assisted, and labs are openly experimenting with systems that propose and run their own experiments. Framing what a result actually means, and taking responsibility for a published claim, have proven much harder to hand over.

How much do AI researchers earn?

It varies enormously by employer and seniority. The US median for the broader 'computer and information research scientist' category was $140,910 in 2024, but new PhDs joining a large technology company's research lab often start well above that, and a small number of senior researchers at frontier AI labs have reportedly been offered total packages worth seven or eight figures during the 2024–2025 hiring competition.

What's the difference between an AI researcher and a machine learning engineer?

The lines blur constantly, but a rough distinction holds: a researcher's output is usually a paper, a new method or an insight about why something works, judged by peer review; an engineer's output is usually a working system in production, judged by whether it ships and holds up under real use. Many people do both across a career.

Which programming languages and tools does AI research use?

Python dominates, paired with a deep-learning framework — PyTorch is now the most common in research, with JAX popular for large-scale and Google-affiliated work. Beyond code, researchers depend on shared GPU or TPU compute clusters, experiment-tracking tools like Weights & Biases, and arXiv, where most new results circulate as preprints months before formal peer review.

Open compare lab

Share this page

AI research is the profession most directly exposed to its own output. The tools it builds are now being turned, deliberately, on the research process itself — summarizing papers, proposing experiments, writing training code — which makes this section less hypothetical than in almost any other profile on this site.

That exposure does not make the job obsolete, but it does make the argument about what survives unusually concrete: a handful of labs are openly trying to build systems that run the entire research loop, and the honest answer, as of now, is that they can do parts of it but not the parts that matter most for judging whether a result is true.

50 / 100
Moderate

Share of the work a machine could do

A meaningful share of an AI researcher's routine work — literature triage, first-draft experiment code, hyperparameter search, drafting a related-work section — is already comparably fast for an AI system to do, and multi-agent research systems built explicitly to automate more of the loop are an active project at several major labs. What has not been automated is judging whether a result is real, deciding which question is worth asking, and being accountable when a published claim turns out to be wrong.

Scored from the tasks, not the job title. Lower is safer.

Jobs AI cannot take →

What machines cannot take

Judging whether a result is real

82

Distinguishing a genuine effect from noise, a lucky seed, or a benchmark quietly leaking into training data requires exactly the skeptical judgment automated systems are worst at applying to their own output.

Choosing which question is worth asking

85

Deciding that a specific, hard, under-explored problem is worth months of a lab's compute and attention is a bet made with incomplete information — the part of research furthest from pattern-matching over past papers.

Accountability for published claims

88

Someone has to stand behind a result when it fails to replicate or turns out to have a flawed baseline — a model cannot hold professional or scientific accountability for its own output.

Cross-disciplinary and ethical judgment

76

Weighing a system's likely social effects, safety risks or dual-use potential before publishing draws on values and context an automated research pipeline has no grounds to reason about on its own.

Building scientific taste over years

78

The intuition for which surprising result is worth chasing, developed by watching hundreds of experiments succeed and fail over a career, is not yet something any system has demonstrated at a senior researcher's level.

What they already take

Literature triage and summarization

70

Given tens of thousands of new preprints a year, tools that summarize and rank papers by relevance are already in routine use across the field.

First-draft experiment and pipeline code

68

Wiring together a standard training loop, data loader or evaluation script is now often faster to generate and check than to write from scratch.

Hyperparameter search

74

Automated tuning systems can search a parameter space far more exhaustively and cheaply than a researcher manually trying configurations one at a time.

Drafting related-work and background sections

55

A first pass at summarizing prior work relevant to a new paper is increasingly AI-assisted, though the specific framing and honest comparison still need a researcher's check.

How the work is changing

From running experiments to directing them

A growing share of a researcher's time goes to specifying what an automated system should try and judging its output, rather than writing and running every experiment personally.

Compute becomes as decisive as ideas

As the field's own 'bitter lesson' predicts, access to large-scale compute increasingly determines which labs can even test a given idea, concentrating frontier research inside a small number of well-funded organizations.

Reproducibility and evaluation gain status

As AI-assisted tools make it easier to produce a plausible-looking result quickly, rigorously verifying that result is becoming a more valued and more separately staffed specialty within research groups.

Automating the research loop becomes its own research topic

Systems designed to propose hypotheses, run experiments and draft papers with limited human oversight — projects like Sakana AI's 'AI Scientist' (2024) and Google DeepMind's 'AI co-scientist' (2025) — are themselves now an active area of publication.

New jobs branching off

AI safety / alignment researcher

Studies how to keep increasingly capable systems reliably doing what their designers intend, a specialty that barely existed as a distinct career path before the 2010s.

ML infrastructure / systems engineer

Builds and maintains the distributed training and compute infrastructure frontier research now depends on as much as any individual algorithmic idea.

AI policy researcher

Works at the boundary between technical research and government or corporate policy, translating what a lab's own systems can and cannot do into rules or safeguards.

Red-teaming / evaluation specialist

Deliberately tests models and AI-generated research outputs for failure modes, security holes and false claims before they reach production or publication.

AI exposure scenarios

Three reversible lenses: augment the work, replace a slice, or open a niche. Teaching marks — not forecasts.

Augment

Keep the role; AI speeds drafts, triage, or research while judgement and accountability stay human.

Replace a slice

A narrow task stack may compress first (templates, first drafts, routine scoring) while adjacent craft grows.

New niche

Oversight, integration, and domain QA roles can appear where AI output must be trusted in regulated settings.

Outlook

Demand for AI research talent is very likely to stay strong through the next decade, but the entry-level version of the job is already changing: fewer roles will exist purely to run routine experiments by hand, and more will expect a new researcher to direct, question and verify work an AI-assisted pipeline produced almost immediately.

The clearest long-run advantage will sit with researchers who can frame a genuinely new question, spot when a plausible-looking result is quietly wrong, and take responsibility for a claim in public — the same judgment the field has always needed, now applied to a research process that increasingly writes its own first draft.

Similar professions

Closest neighbours on the six-score profile — not the same field only.

Continue exploring

Keep exploring

More in Engineering & Technology