Skip to content Skip to a section

📊The Greats

Data Scientist · Finds patterns and builds predictive models from data — a 2008 job title built on three centuries of counting, testing and visualizing evidence.

At a glance
Score intensity

Darker cells mean a higher score for this topic on that metric.

Last reviewed Sources & creditsMedia creditsMethodology

Quick answers

What does a data scientist actually do day to day?

Contrary to the 'building AI models' image, most days split between writing SQL queries and Python code to pull and clean data, running statistical tests or training models, and translating results into a chart or memo a non-technical stakeholder can act on. Practitioner surveys consistently put data cleaning and preparation at roughly half or more of total working time.

Do you need a PhD to become a data scientist?

No. Unlike medicine or law, there is no license or single required degree; a bachelor's or master's in statistics, computer science, mathematics or a related quantitative field is the most common path, and a strong portfolio of real projects often matters more to employers than the exact credential. PhDs are more common in research-heavy or applied machine-learning roles.

Is data science just statistics with a new name?

Not quite. It draws heavily on statistics — Fisher's experimental design, Tukey's exploratory analysis — but adds programming, database engineering and machine learning that classical statistics departments rarely taught. William S. Cleveland proposed the term in 2001 specifically to describe this enlarged, computing-heavy version of the field, built on statistics rather than replacing it.

Is data science at risk from AI?

The routine end is already exposed: AutoML tools can fit and tune standard models, and AI assistants can write SQL queries, first-draft exploratory charts and boilerplate pipeline code faster than a person. What has not been automated is framing the right question, judging whether a pattern is meaningful or spurious, and taking responsibility for a decision built on the result.

How much do data scientists earn?

It varies widely by country and seniority. The US Bureau of Labor Statistics put the median annual wage for the Data Scientists occupation at roughly $108,000 in 2023, while junior analysts often start closer to $70,000-$95,000 and senior or principal data scientists at major technology companies can earn $200,000-$400,000 or more in total compensation with equity.

What's the difference between a data scientist, a data analyst and a machine learning engineer?

A data analyst typically answers defined business questions with existing data and dashboards; a data scientist builds new statistical models and predictive analyses, often from messier data; a machine learning engineer takes a model out of a notebook and makes it run reliably, at scale, in production. The lines blur constantly, and many people move between all three across a career.

Open compare lab

Share this page

Data science's canon runs from a 17th-century London tradesman counting plague deaths by hand to a 21st-century scientist mining billions of anonymized link clicks. The eight people here span more than three centuries and three continents.

What connects them is less a shared method than a shared instinct: look hard at real, messy records, and trust a pattern only after checking it survives scrutiny — the same discipline the field still runs on, however it happens to be titled this decade.

The all-time podium

Florence Nightingale
Florence Nightingale
United Kingdom
2
Ronald Fisher
Ronald Fisher
United Kingdom
1
William Sealy Gosset
William Sealy Gosset
United Kingdom / Ireland
3

“A rigorous experiment needs a control and randomization built in from the start, not argued about afterward.”

Ronald Fisher

The eight who reached the top

1
Photograph of Ronald Fisher Unknown author Unknown author · Public domain

Ronald Fisher

United Kingdom · 1890–1962

British statistician and geneticist who built much of the modern framework of applied statistics — analysis of variance, randomized experimental design and maximum likelihood estimation — largely while analyzing decades of crop-yield trials at Rothamsted Experimental Station in the 1920s.

The story

At a Rothamsted afternoon tea in the late 1920s, colleague Muriel Bristol claimed she could tell whether milk had been poured into a cup before or after the tea. Rather than dismiss her, Fisher designed a randomized experiment to actually test the claim — the scenario he later used in his 1935 book 'The Design of Experiments' to lay out the logic of the null hypothesis and randomized trial design.

“A rigorous experiment needs a control and randomization built in from the start, not argued about afterward.”

'Statistical Methods for Research Workers' published
1925
Years at Rothamsted Experimental Station
~14 yrs
Royal Society Fellow elected
1929
2
Photograph of Florence Nightingale Henry Hering (1814-1893) · Public domain

Florence Nightingale

United Kingdom · 1820–1910

Better known as the founder of modern nursing, Nightingale was also a pioneering statistician who used data visualization to convince the British government that poor sanitation, not battle wounds, was killing most soldiers in the Crimean War, and became the first woman elected a Fellow of the Royal Statistical Society.

The story

In 1858, Nightingale published her 'coxcomb' or polar-area diagram — a circular chart dividing the year into wedges sized by monthly deaths, colored by cause — to show the War Office that preventable disease, not combat, caused the overwhelming majority of British Army deaths in the Crimea; the visualization helped drive lasting army sanitary reforms.

“A well-designed chart can change government policy faster than the same numbers in a table ever will.”

Elected Fellow, Royal Statistical Society
1858 (first woman)
'Notes on Matters Affecting Health' published
1858
Royal Sanitary Commission established
1857
3
Photograph of William Sealy Gosset User Wujaszek on pl.wikipedia · Public domain

William Sealy Gosset

United Kingdom / Ireland · 1876–1937

Working as a chemist and statistician for the Guinness brewery in Dublin, Gosset developed the t-distribution and t-test to solve a very practical problem — judging the quality of small barley samples — and published under the pseudonym 'Student' because Guinness barred employees from publishing under their own names.

The story

Guinness's concern about competitors learning trade secrets led it to forbid staff publications; Gosset's landmark 1908 paper 'The Probable Error of a Mean' therefore appeared in Biometrika under the pseudonym 'Student,' a name still attached to the t-test used in millions of small-sample analyses today, long after his true identity became known among statisticians.

“A method built to solve one narrow, practical problem can outlive its original purpose by more than a century.”

'The Probable Error of a Mean' published
1908, Biometrika
Years at Guinness
~38 yrs
Pseudonym used in publications
'Student'
4
Photograph of C. R. Rao Prateek Karandikar · CC BY-SA 4.0

C. R. Rao

India / United States · 1920–2023

Indian-American statistician who, at age 25, derived the Cramér–Rao bound and the Rao–Blackwell theorem — two of the most widely used results in statistical estimation theory — and later helped build India's Indian Statistical Institute before a long career at American universities.

The story

Rao's 1945 paper, written as a young researcher at the Indian Statistical Institute in Kolkata, proved a lower limit on how precisely any unbiased estimator can possibly measure an unknown quantity — now called the Cramér–Rao bound — a result so foundational it appears in nearly every statistics textbook published since, credited to a paper he wrote before turning 26.

“A single, precisely proven result early in a career can outlast decades of a field's later fashions.”

Cramér–Rao bound paper published
1945
US National Medal of Science
2002
Age at death
102
5

John Tukey

United States · 1915–2000

American mathematician and statistician who co-invented the Fast Fourier Transform algorithm, invented the box plot, and argued in his 1962 paper 'The Future of Data Analysis' that analyzing real data should be recognized as its own scientific discipline distinct from mathematical statistics.

The story

In 'The Future of Data Analysis' (1962), Tukey wrote that data analysis had been 'ill-taught' as a minor branch of mathematics, and spent the next fifteen years developing hands-on techniques — including the box plot and stem-and-leaf plot, introduced in his 1977 book 'Exploratory Data Analysis' — designed to be worked by hand with pencil and paper before a computer ever touched the numbers.

“Look at the data with your own eyes, using simple hand-drawn tools, before trusting any model built from it.”

'The Future of Data Analysis' published
1962
'Exploratory Data Analysis' published
1977
Fast Fourier Transform algorithm
1965, with Cooley
6

William S. Cleveland

United States · b. 1943

Bell Labs statistician who coined 'data science' as the name for an enlarged statistics discipline in a 2001 paper, and separately invented widely used data-visualization and smoothing techniques including the LOESS local-regression method.

The story

Cleveland's 2001 paper 'Data Science: An Action Plan for Expanding the Technical Areas of the Field of Statistics' proposed that universities create dedicated data-science departments distinct from statistics departments — a specific institutional recommendation that took over a decade to catch on, but by the mid-2010s hundreds of universities worldwide had done exactly that.

“Naming and formally proposing a field's boundaries in print can shape how universities organize around it years later.”

'Data Science: An Action Plan' published
2001
'Visualizing Data' published
1993
Years at Bell Labs
~30 yrs
7

DJ Patil

United States · b. 1973

Led LinkedIn's data team where he and colleague Jeff Hammerbacher independently settled on the job title 'data scientist' around 2008, co-wrote the influential 2012 Harvard Business Review article naming it 'the sexiest job of the 21st century,' and later served as the first US Chief Data Scientist under President Obama.

The story

Patil has recounted that while building LinkedIn's data team around 2008, he and his team considered titles like 'data analyst' and 'business analyst' before settling on 'data scientist' partly because it better captured the mix of engineering and statistics the role actually required, and partly because it was simply a more appealing title to recruit for.

“The right job title can shape how a whole profession is recruited into, almost as much as the work itself.”

'Sexiest Job' HBR article published
Oct 2012, with Davenport
US Chief Data Scientist tenure
2015–2017
LinkedIn data team joined
2008
8

Hilary Mason

United States · b. 1978

Became bit.ly's chief scientist in 2009, one of the earliest people to hold the job title publicly, using the link-shortening service's massive click-stream data to study what spreads online, and later founded the machine-learning research startup Fast Forward Labs.

The story

At bit.ly, Mason and colleagues mined billions of anonymized shortened-link clicks to study how information actually spreads across the internet in real time, publishing findings on patterns like which time of day links get clicked most — turning a link-shortening utility into one of the first large-scale, publicly discussed data science research programs.

“A mundane infrastructure product can double as one of the richest research datasets available, if someone thinks to look.”

bit.ly Chief Scientist since
2009
Fast Forward Labs founded
2014
Fast Forward Labs acquired by Cloudera
2017

Bars are scaled to the leader in this list.

Comparison Lab

Toggle names on and off — every bar rescales to the leader of your selection.

5 / 8

The argument

Whether DJ Patil and Jeff Hammerbacher truly 'coined' the job title is disputed; both have credited broader teams and informal prior industry use, and the phrase appears in scattered references before 2008.

Credit for naming the field itself is contested: computer scientist Peter Naur used the term 'data science' as a synonym for computer science as early as 1974, decades before William S. Cleveland's widely cited 2001 proposal.

Similar professions

Closest neighbours on the six-score profile — not the same field only.

Continue exploring

Keep exploring

More in Science & Research