📊Origins & Evolution

Data Scientist · Finds patterns and builds predictive models from data — a 2008 job title built on three centuries of counting, testing and visualizing evidence.

Data science is one of the youngest job titles on this site sitting on top of one of the oldest crafts: turning a pile of records into a claim someone can act on. The name 'data scientist' dates only to around 2008; the underlying discipline runs back through three centuries of statisticians counting, testing and visualizing evidence.

Two threads run through the whole story: a slow professionalization of statistical inference, from a London tradesman's mortality tables to Fisher's crop trials, and a much faster, more recent packaging of that inheritance into a corporate job title, an academic department and, briefly, the 'sexiest job of the 21st century.'

Where it began

1662London, England

London haberdasher John Graunt published 'Natural and Political Observations ... Made upon the Bills of Mortality,' a systematic analysis of the weekly parish death records London had kept since the plague years, from which he estimated the city's population and identified patterns in causes of death. It was one of the earliest known attempts to draw general conclusions from a large body of collected data rather than individual anecdote, and it earned Graunt, a shopkeeper with no university training, election to the Royal Society — the intellectual root the field's much later name would eventually be attached to.

Timeline

1662Graunt analyzes London's Bills of Mortality

John Graunt publishes a statistical study of London's parish death records, estimating population size and mortality patterns — among the earliest systematic uses of collected data to draw general conclusions.

1908Gosset publishes the t-test as 'Student'

William Sealy Gosset, a statistician for the Guinness brewery in Dublin, publishes 'The Probable Error of a Mean' under the pseudonym 'Student' because Guinness barred staff from publishing under their own names, introducing the t-distribution used in small-sample analysis ever since.

1925Fisher publishes 'Statistical Methods for Research Workers'

Ronald Fisher, analyzing decades of crop-yield data at Rothamsted Experimental Station, publishes the book that establishes analysis of variance and modern experimental design as standard statistical practice.

1933Neyman and Pearson formalize hypothesis testing

Jerzy Neyman and Egon Pearson's paper 'On the Problem of the Most Efficient Tests of Statistical Hypotheses' gives statistics its modern framework for deciding, with quantified error rates, whether a result is significant.

1962Tukey argues data analysis is its own science

John Tukey's paper 'The Future of Data Analysis' argues that analyzing real data deserves recognition as a distinct science with its own methods, not a minor branch of mathematical statistics.

1977Tukey's 'Exploratory Data Analysis' popularizes the box plot

Tukey's book gives practitioners hands-on tools — the box plot among them — for understanding a dataset's shape by eye before fitting any model to it.

1996Fayyad, Piatetsky-Shapiro and Smyth define 'knowledge discovery'

The trio's paper in AI Magazine formalizes 'Knowledge Discovery in Databases' as the broader process surrounding raw data mining, part of a 1990s wave of formalizing pattern-discovery in growing corporate databases.

2001Cleveland proposes 'data science' as a field

Bell Labs statistician William S. Cleveland publishes 'Data Science: An Action Plan for Expanding the Technical Areas of the Field of Statistics,' proposing dedicated university data-science departments.

2008Patil and Hammerbacher coin the job title

Building data teams at LinkedIn and Facebook respectively, DJ Patil and Jeff Hammerbacher independently settle on 'data scientist' as a job title for the statistics-plus-engineering work their teams were doing.

2012Harvard Business Review calls it the 'sexiest job'

Thomas Davenport and DJ Patil's October 2012 article 'Data Scientist: The Sexiest Job of the 21st Century' triggers a decade-long corporate hiring boom and a wave of new university degree programs.

The eras

1662–1899

Counting the living and the dead

In 1662, London haberdasher John Graunt published a statistical analysis of the city's parish death records, one of the earliest known attempts to draw general conclusions from collected data rather than individual anecdote. Belgian astronomer Adolphe Quetelet extended the idea in the 1830s with his 'social physics,' applying statistical distributions to human traits. Florence Nightingale carried the tradition into practical persuasion in 1858, using a polar-area diagram of Crimean War mortality data to convince the British War Office that poor sanitation, not combat, was killing most soldiers — becoming the first woman elected a Fellow of the Royal Statistical Society in the process.

Photograph of Ronald Fisher, whose statistical methods underlie modern data science
Unknown author Unknown author · Public domain · Wikimedia Commons
1900–1935

Statisticians formalize inference

Karl Pearson founded the journal Biometrika in 1901 to give the young discipline of mathematical statistics a home. William Sealy Gosset, working as a statistician for the Guinness brewery, published the t-test in 1908 under the pseudonym 'Student' because Guinness barred employees from publishing under their own names. Ronald Fisher's 1925 book 'Statistical Methods for Research Workers,' written while analyzing decades of crop trials at Rothamsted Experimental Station, introduced analysis of variance and rigorous experimental design. Jerzy Neyman and Egon Pearson completed the toolkit in 1933 with a formal framework for hypothesis testing — the statistical machinery data science would later inherit almost intact.

ENIAC, one of the first electronic general-purpose computers, unveiled in 1946
The original uploader was TexasDex at English Wikipedia . · CC BY-SA 3.0 · Wikimedia Commons
1936–1969

Computers arrive, and Tukey makes a case

World War II accelerated large-scale calculation, from human computing teams to the first electronic general-purpose computers such as ENIAC, unveiled in 1946. Statisticians who had spent decades on hand calculation increasingly used machines to run analyses at a scale previously impossible. In 1962, Bell Labs and Princeton statistician John Tukey published 'The Future of Data Analysis,' arguing that analyzing real data was being 'ill-taught' as a minor branch of mathematical statistics and deserved recognition as its own science, with its own methods and its own claim on universities' attention.

A box plot, one of the exploratory data analysis tools John Tukey popularized
User:Schutz · Public domain · Wikimedia Commons
1970–1999

Exploratory analysis and the data-mining boom

Tukey's 1977 book 'Exploratory Data Analysis' gave practitioners hands-on tools — the box plot among them — for understanding a dataset's shape before fitting any model to it. As businesses accumulated growing transactional databases through the 1980s and 1990s, a parallel discipline of 'data mining' emerged to find patterns in them; a 1996 paper by Usama Fayyad, Gregory Piatetsky-Shapiro and Padhraic Smyth formalized the distinction between raw data mining and the broader process of 'Knowledge Discovery in Databases.' Corporate data warehousing, popularized through consultant Bill Inmon's writing across the decade, gave this work a permanent home inside large companies.

Conceptual illustration of large-scale data, the raw material of 21st-century data science
Jouasse · CC BY-SA 4.0 · Wikimedia Commons
2000–present

Named, titled, hyped, and specializing

Bell Labs statistician William S. Cleveland proposed 'data science' as the name for an enlarged statistics discipline in a 2001 paper. Google's 2004 MapReduce paper and the open-source Hadoop framework that followed gave the field cheap tools for internet-scale data. Around 2008, DJ Patil and Jeff Hammerbacher, building data teams at LinkedIn and Facebook, independently settled on 'data scientist' as a job title, and Thomas Davenport and Patil's October 2012 Harvard Business Review article calling it 'the sexiest job of the 21st century' triggered a decade-long corporate hiring boom. By the late 2010s the role had matured enough to split into more specialized titles — analytics engineer, machine learning engineer, product analyst — even as generative AI began automating much of the routine querying and charting the original job description assumed a person would do by hand.

What this job replaced

Neighbouring trades that no longer exist — absorbed, automated or regulated away.

Human computer

1880s–1950s

Long before electronic computers, teams of workers — disproportionately women, such as the 'Harvard Computers' who catalogued stars at the Harvard College Observatory from the 1880s, and NASA's Katherine Johnson, Dorothy Vaughan and Mary Jackson, who hand-calculated orbital trajectories into the 1960s — performed the mathematical labor now done by software. Electronic computers absorbed the role within a generation of becoming reliable and affordable.

Punch-card / keypunch operator

1890s–1970s

Following Herman Hollerith's punched-card tabulator, built for the 1890 US Census, keypunch operators spent decades transcribing paper records onto cards that mainframe computers could read, a specialized clerical skill taught in dedicated training courses. Direct-entry computer terminals made the role largely obsolete by the 1970s and 1980s.

Mainframe batch-report programmer

1960s–2000s

Corporate 'data processing' departments once employed programmers, often working in COBOL, whose main job was writing and running scheduled overnight batch jobs that produced management's printed reports by the next morning. Self-service business-intelligence dashboards and ad hoc query tools let managers pull their own numbers on demand, closing this specialty out almost entirely by the 2000s.

Trades that vanished →

Every capability data science claims as new — pattern-finding in large datasets, rigorous inference, persuasive visualization — has a specific, checkable ancestor: Graunt's mortality tables, Fisher's crop trials, Tukey's box plots. What changed in 2001 and 2008 was mostly the name and the packaging, applied to an old craft newly powered by cheap storage and computing.

The 2012 'sexiest job' boom that followed has already partly run its course: the title has splintered into analytics engineer, machine learning engineer and product analyst, and generative AI is absorbing the routine querying and charting the original 2008 job description took for granted. The underlying discipline — deciding what a pattern in the data actually means — looks likely to outlast the specific job title once again.

Keep exploring

More in Science & Research