Look at the data before you model it
Plot the raw distributions, scatterplots and simple summaries before fitting anything — a habit that catches broken data, unit errors and outliers a model would otherwise quietly absorb and hide.
Data Scientist · Finds patterns and builds predictive models from data — a 2008 job title built on three centuries of counting, testing and visualizing evidence.
Darker cells mean a higher score for this topic on that metric.
LessMore
Last reviewed Sources & creditsMedia creditsMethodology
Contrary to the 'building AI models' image, most days split between writing SQL queries and Python code to pull and clean data, running statistical tests or training models, and translating results into a chart or memo a non-technical stakeholder can act on. Practitioner surveys consistently put data cleaning and preparation at roughly half or more of total working time.
No. Unlike medicine or law, there is no license or single required degree; a bachelor's or master's in statistics, computer science, mathematics or a related quantitative field is the most common path, and a strong portfolio of real projects often matters more to employers than the exact credential. PhDs are more common in research-heavy or applied machine-learning roles.
Not quite. It draws heavily on statistics — Fisher's experimental design, Tukey's exploratory analysis — but adds programming, database engineering and machine learning that classical statistics departments rarely taught. William S. Cleveland proposed the term in 2001 specifically to describe this enlarged, computing-heavy version of the field, built on statistics rather than replacing it.
The routine end is already exposed: AutoML tools can fit and tune standard models, and AI assistants can write SQL queries, first-draft exploratory charts and boilerplate pipeline code faster than a person. What has not been automated is framing the right question, judging whether a pattern is meaningful or spurious, and taking responsibility for a decision built on the result.
It varies widely by country and seniority. The US Bureau of Labor Statistics put the median annual wage for the Data Scientists occupation at roughly $108,000 in 2023, while junior analysts often start closer to $70,000-$95,000 and senior or principal data scientists at major technology companies can earn $200,000-$400,000 or more in total compensation with equity.
A data analyst typically answers defined business questions with existing data and dashboards; a data scientist builds new statistical models and predictive analyses, often from messier data; a machine learning engineer takes a model out of a notebook and makes it run reliably, at scale, in production. The lines blur constantly, and many people move between all three across a career.
The popular image of a data scientist building clever predictive models is real but incomplete. Practitioner surveys and firsthand accounts consistently describe a job where finding, cleaning and reshaping messy data eats up more hours than the modeling step most people picture.
The craft passed down inside the field is less about memorizing a specific algorithm and more about disciplined skepticism: looking at the data before trusting any model of it, checking whether a pattern holds inside subgroups, and knowing when the honest answer to a stakeholder's question is that the data cannot support one.
Hypothesis testing, regression, and understanding what a p-value or a confidence interval actually claims — the inferential foundation everything else in the job sits on top of.
Writing code fluent enough to pull, clean and transform data at scale, not just prototype an analysis in a notebook that never has to run again.
Turning messy, inconsistent, real-world data — missing values, duplicate records, mismatched formats — into something a model or chart can actually use.
Choosing, training and validating a model appropriate to the problem, and knowing when a simple regression beats an elaborate neural network.
Translating a statistical result into a chart, a memo or a recommendation a non-technical executive can act on, without oversimplifying or burying the real uncertainty.
Understanding what a metric actually means inside a specific business or scientific field, so an analysis answers the question that was actually asked.
Reviewing overnight pipeline runs and key metrics dashboards, and a short sync with the team on priorities for the day.
Writing SQL queries against a data warehouse, joining and reshaping tables, and handling missing or inconsistent values — commonly the largest single block of a working day.
A genuine break, and often an informal setting where cross-team problems get raised before they reach a scheduled meeting.
Building or refining a statistical model, running an experiment's significance test, or digging into why a metric moved.
Presenting findings to a product or business team, reviewing an A/B test's results, or defending an analysis's assumptions under questioning.
Personal time, except around a product launch or a quarterly business review, when evenings can absorb last-minute analysis requests.
Craft knowledge practitioners actually pass on — not motivation.
Plot the raw distributions, scatterplots and simple summaries before fitting anything — a habit that catches broken data, unit errors and outliers a model would otherwise quietly absorb and hide.
An aggregate trend can reverse completely once the data is split by a relevant category — a pattern known as Simpson's paradox, named for a 1951 paper by statistician Edward H. Simpson.
Repeatedly checking performance against the same held-out data, then adjusting the model, quietly leaks information from that data into the model — touching it only once is what keeps a final result honest.
Real projects consistently run closer to 80% data preparation and 20% modeling than the reverse — planning a timeline around the opposite assumption is a common way projects run late.
A baseline — sometimes just a rule or a linear model — clarifies how much a more complex model actually adds, and is often good enough to ship on its own.
A surprising correlation is more often an artifact of a hidden third variable than a genuine discovery — actively search for what else could explain it before presenting a result.
The dominant language for data manipulation and classical machine learning, with pandas for tabular data and scikit-learn for standard modeling algorithms.
The query language for pulling data out of the relational and cloud data warehouses — Snowflake, BigQuery, Redshift — where most companies' data actually lives.
The standard interactive environment for exploratory analysis, letting a data scientist run code, view a chart and take notes in the same document.
Business-intelligence dashboarding tools used to turn a finished analysis into something a non-technical stakeholder can explore and monitor without writing code.
Version control for code and, increasingly, for models, alongside the cloud infrastructure that stores data and runs training jobs at a scale a laptop cannot handle.
Running many statistical comparisons and reporting only the one that crosses a significance threshold produces false discoveries that will not replicate — a well-documented failure mode known as p-hacking.
Accidentally including a feature that encodes the outcome being predicted produces unrealistically strong offline results that collapse once the model runs on genuinely unseen production data.
A carefully tuned machine-learning model solving a problem a simple SQL query or spreadsheet formula would have answered just as well wastes effort and can miss the actual business question.
Closest neighbours on the six-score profile — not the same field only.
The keeper of the books: heir to a craft so old it invented writing itself, now negotiating with the software built to automate it.
AI-resistant 35 🗺️Decides what a company should build next, and why — turning customer needs, business goals and engineering limits into one shared plan nobody else fully owns.
AI-resistant 50 💻Writes, tests and maintains the code that runs modern life — and is one of the first professions watching AI automate its own daily work.
AI-resistant 35 📣The professional who creates demand — from Pompeii's painted walls and P&G's 1931 brand-man memo to the auction-driven feeds of the digital era.
AI-resistant 38 📦Builds the systems that train, deploy, monitor and govern machine-learning models in production.
AI-resistant 54 🔎Studies how people use products, turning observed behavior, needs and frustrations into evidence teams can design around.
AI-resistant 69Derives and tests the mathematical laws governing matter, energy, space and time, from a lone chalkboard to a 3,000-author particle-collider paper.
AI-resistant 65 🧬The scientist who studies life itself, from Linnaeus naming species by hand to editing genomes with CRISPR, still testing every idea against a living organism.
AI-resistant 64 🧪The scientist who makes and measures matter itself — from Tapputi's Babylonian perfume still to today's robot laboratories, still the one who decides what the spectrum means.
AI-resistant 70 🔭The scientist who measures the universe, from Babylonian clay tablets to space telescopes, still deciding which flicker in the data is a discovery.
AI-resistant 70 🧮From Babylonian scribes to Fields medalists and AI-assisted proof: the profession that turns hard questions into permanent certainty, one theorem at a time.
AI-resistant 70 🧿Builds, measures and controls devices that exploit quantum states for computing, sensing, communication and materials research.
AI-resistant 78 🧫Uses clinical, trial and health-system data to generate reliable evidence for safer care, research and operational decisions.
AI-resistant 68