Framing the right question
84Turning a vague business or scientific problem into a specific, testable question requires organizational and domain context an automated tool has no access to on its own.
Data Scientist · Finds patterns and builds predictive models from data — a 2008 job title built on three centuries of counting, testing and visualizing evidence.
Darker cells mean a higher score for this topic on that metric.
LessMore
Last reviewed Sources & creditsMedia creditsMethodology
Contrary to the 'building AI models' image, most days split between writing SQL queries and Python code to pull and clean data, running statistical tests or training models, and translating results into a chart or memo a non-technical stakeholder can act on. Practitioner surveys consistently put data cleaning and preparation at roughly half or more of total working time.
No. Unlike medicine or law, there is no license or single required degree; a bachelor's or master's in statistics, computer science, mathematics or a related quantitative field is the most common path, and a strong portfolio of real projects often matters more to employers than the exact credential. PhDs are more common in research-heavy or applied machine-learning roles.
Not quite. It draws heavily on statistics — Fisher's experimental design, Tukey's exploratory analysis — but adds programming, database engineering and machine learning that classical statistics departments rarely taught. William S. Cleveland proposed the term in 2001 specifically to describe this enlarged, computing-heavy version of the field, built on statistics rather than replacing it.
The routine end is already exposed: AutoML tools can fit and tune standard models, and AI assistants can write SQL queries, first-draft exploratory charts and boilerplate pipeline code faster than a person. What has not been automated is framing the right question, judging whether a pattern is meaningful or spurious, and taking responsibility for a decision built on the result.
It varies widely by country and seniority. The US Bureau of Labor Statistics put the median annual wage for the Data Scientists occupation at roughly $108,000 in 2023, while junior analysts often start closer to $70,000-$95,000 and senior or principal data scientists at major technology companies can earn $200,000-$400,000 or more in total compensation with equity.
A data analyst typically answers defined business questions with existing data and dashboards; a data scientist builds new statistical models and predictive analyses, often from messier data; a machine learning engineer takes a model out of a notebook and makes it run reliably, at scale, in production. The lines blur constantly, and many people move between all three across a career.
Data science sits at an unusual point in the AI story: many of the tools reshaping the profession — automated machine learning, natural-language-to-SQL translation, AI-assisted charting — were themselves largely built by data scientists and machine-learning researchers, now aimed back at parts of their own daily work.
That exposure is real and uneven. The routine, well-specified end of the job — pulling data, running a standard model, producing a first-draft chart — is already comparably fast for AI tools to do. The end that survives is less about writing code and more about deciding what question is worth asking and whether an answer can be trusted.
A meaningful share of routine data-science work — SQL querying, exploratory charting, fitting a standard model with AutoML, writing boilerplate pipeline code — is already automatable or AI-assisted today, and that share is growing quickly. What resists automation is framing an ambiguous business problem as a testable question, judging whether a result is real or a statistical artifact, and being accountable when a model's real-world decision turns out to be wrong.
Scored from the tasks, not the job title. Lower is safer.
Jobs AI cannot take →Turning a vague business or scientific problem into a specific, testable question requires organizational and domain context an automated tool has no access to on its own.
Telling a genuine effect apart from noise, a data-leakage artifact or a lucky train-test split is exactly the skeptical judgment automated pipelines are worst at applying to their own output.
Someone has to answer for a model that denies a loan, flags a patient or misprices a product — a responsibility that cannot transfer to the tool that produced the number.
Getting a counterintuitive or unwelcome finding actually acted on depends on trust, credibility and negotiation with people, not just a correct analysis.
Choosing the right natural experiment, control group or causal-inference strategy for a new question is closer to creative judgment than to pattern-matching over past analyses.
Natural-language-to-SQL tools now handle a large share of straightforward data-retrieval requests that once required a data scientist's direct involvement.
Automated exploratory-analysis tools can generate distribution plots, correlation summaries and basic charts from a raw dataset in seconds.
AutoML platforms can select, train and tune conventional models like gradient boosting or logistic regression largely unattended, for well-defined, well-labeled problems.
Wiring together a standard data-transformation or loading script is now often faster to generate with an AI coding assistant and review than to write from scratch.
A growing share of the job is specifying what an automated pipeline should try and verifying its output, rather than writing every query and model by hand.
Analytics engineering, machine learning engineering and product analytics have increasingly become separate job titles, each absorbing a slice of what a single 'data scientist' was expected to cover a decade ago.
As pattern-finding in existing data becomes easier to automate, the harder skill of designing a trustworthy experiment or causal analysis is becoming more valued, not less.
No-code and AutoML tools let non-specialists run their own basic analyses, pushing dedicated data scientists toward the harder, more ambiguous problems those tools cannot handle.
Builds and maintains the data-transformation pipelines — often using tools like dbt — that feed dashboards and models, sitting between data engineering and analysis.
Takes a model out of a research notebook and makes it run reliably, monitored and at scale in a production system.
Owns the roadmap for a data or AI-driven feature, translating business goals into what a data team should actually build.
Audits models and the data behind them for bias, fairness and regulatory compliance before they are allowed to ship.
Three reversible lenses: augment the work, replace a slice, or open a niche. Teaching marks — not forecasts.
Keep the role; AI speeds drafts, triage, or research while judgement and accountability stay human.
A narrow task stack may compress first (templates, first drafts, routine scoring) while adjacent craft grows.
Oversight, integration, and domain QA roles can appear where AI output must be trusted in regulated settings.
Demand for people who can turn data into a trustworthy decision is very likely to stay strong through the next decade, but the entry-level version of the job is already changing: fewer roles will exist purely to write routine queries and first-draft charts by hand.
The clearest long-run advantage sits with data scientists who can frame a genuinely useful question, spot when an AI-generated result is quietly wrong, and take responsibility for a recommendation in front of skeptical stakeholders — the same judgment the field has needed since Fisher's crop trials, now applied to a workflow that increasingly drafts its own first pass.
Closest neighbours on the six-score profile — not the same field only.
The keeper of the books: heir to a craft so old it invented writing itself, now negotiating with the software built to automate it.
AI-resistant 35 🗺️Decides what a company should build next, and why — turning customer needs, business goals and engineering limits into one shared plan nobody else fully owns.
AI-resistant 50 💻Writes, tests and maintains the code that runs modern life — and is one of the first professions watching AI automate its own daily work.
AI-resistant 35 📣The professional who creates demand — from Pompeii's painted walls and P&G's 1931 brand-man memo to the auction-driven feeds of the digital era.
AI-resistant 38 📦Builds the systems that train, deploy, monitor and govern machine-learning models in production.
AI-resistant 54 🔎Studies how people use products, turning observed behavior, needs and frustrations into evidence teams can design around.
AI-resistant 69Derives and tests the mathematical laws governing matter, energy, space and time, from a lone chalkboard to a 3,000-author particle-collider paper.
AI-resistant 65 🧬The scientist who studies life itself, from Linnaeus naming species by hand to editing genomes with CRISPR, still testing every idea against a living organism.
AI-resistant 64 🧪The scientist who makes and measures matter itself — from Tapputi's Babylonian perfume still to today's robot laboratories, still the one who decides what the spectrum means.
AI-resistant 70 🔭The scientist who measures the universe, from Babylonian clay tablets to space telescopes, still deciding which flicker in the data is a discovery.
AI-resistant 70 🧮From Babylonian scribes to Fields medalists and AI-assisted proof: the profession that turns hard questions into permanent certainty, one theorem at a time.
AI-resistant 70 🧿Builds, measures and controls devices that exploit quantum states for computing, sensing, communication and materials research.
AI-resistant 78 🧫Uses clinical, trial and health-system data to generate reliable evidence for safer care, research and operational decisions.
AI-resistant 68