📊AI & The Future

Data Scientist · Finds patterns and builds predictive models from data — a 2008 job title built on three centuries of counting, testing and visualizing evidence.

Data science sits at an unusual point in the AI story: many of the tools reshaping the profession — automated machine learning, natural-language-to-SQL translation, AI-assisted charting — were themselves largely built by data scientists and machine-learning researchers, now aimed back at parts of their own daily work.

That exposure is real and uneven. The routine, well-specified end of the job — pulling data, running a standard model, producing a first-draft chart — is already comparably fast for AI tools to do. The end that survives is less about writing code and more about deciding what question is worth asking and whether an answer can be trusted.

62 / 100
High

Share of the work a machine could do

A meaningful share of routine data-science work — SQL querying, exploratory charting, fitting a standard model with AutoML, writing boilerplate pipeline code — is already automatable or AI-assisted today, and that share is growing quickly. What resists automation is framing an ambiguous business problem as a testable question, judging whether a result is real or a statistical artifact, and being accountable when a model's real-world decision turns out to be wrong.

Scored from the tasks, not the job title. Lower is safer.

Jobs AI cannot take →

What machines cannot take

Framing the right question

84

Turning a vague business or scientific problem into a specific, testable question requires organizational and domain context an automated tool has no access to on its own.

Judging whether a pattern is real

80

Telling a genuine effect apart from noise, a data-leakage artifact or a lucky train-test split is exactly the skeptical judgment automated pipelines are worst at applying to their own output.

Accountability for a real-world decision

88

Someone has to answer for a model that denies a loan, flags a patient or misprices a product — a responsibility that cannot transfer to the tool that produced the number.

Persuading skeptical stakeholders

68

Getting a counterintuitive or unwelcome finding actually acted on depends on trust, credibility and negotiation with people, not just a correct analysis.

Designing a genuinely new experiment

74

Choosing the right natural experiment, control group or causal-inference strategy for a new question is closer to creative judgment than to pattern-matching over past analyses.

What they already take

Routine SQL queries and data pulls

75

Natural-language-to-SQL tools now handle a large share of straightforward data-retrieval requests that once required a data scientist's direct involvement.

First-draft exploratory charts and summaries

65

Automated exploratory-analysis tools can generate distribution plots, correlation summaries and basic charts from a raw dataset in seconds.

Fitting and tuning standard models

68

AutoML platforms can select, train and tune conventional models like gradient boosting or logistic regression largely unattended, for well-defined, well-labeled problems.

Boilerplate pipeline and ETL code

60

Wiring together a standard data-transformation or loading script is now often faster to generate with an AI coding assistant and review than to write from scratch.

How the work is changing

From building models to directing and checking them

A growing share of the job is specifying what an automated pipeline should try and verifying its output, rather than writing every query and model by hand.

The generalist title splits into specialties

Analytics engineering, machine learning engineering and product analytics have increasingly become separate job titles, each absorbing a slice of what a single 'data scientist' was expected to cover a decade ago.

Causal inference and experimentation gain status

As pattern-finding in existing data becomes easier to automate, the harder skill of designing a trustworthy experiment or causal analysis is becoming more valued, not less.

The 'citizen data scientist' spreads basic analysis

No-code and AutoML tools let non-specialists run their own basic analyses, pushing dedicated data scientists toward the harder, more ambiguous problems those tools cannot handle.

New jobs branching off

Analytics engineer

Builds and maintains the data-transformation pipelines — often using tools like dbt — that feed dashboards and models, sitting between data engineering and analysis.

Machine learning engineer

Takes a model out of a research notebook and makes it run reliably, monitored and at scale in a production system.

Data product manager

Owns the roadmap for a data or AI-driven feature, translating business goals into what a data team should actually build.

Responsible AI / model-risk analyst

Audits models and the data behind them for bias, fairness and regulatory compliance before they are allowed to ship.

Outlook

Demand for people who can turn data into a trustworthy decision is very likely to stay strong through the next decade, but the entry-level version of the job is already changing: fewer roles will exist purely to write routine queries and first-draft charts by hand.

The clearest long-run advantage sits with data scientists who can frame a genuinely useful question, spot when an AI-generated result is quietly wrong, and take responsibility for a recommendation in front of skeptical stakeholders — the same judgment the field has needed since Fisher's crop trials, now applied to a workflow that increasingly drafts its own first pass.

Keep exploring

More in Science & Research