Skip to content

📦The Greats

MLOps Engineer · Builds the systems that train, deploy, monitor and govern machine-learning models in production.

Share this page

MLOps has no single founder because it is an intersection of several disciplines. Its intellectual ancestors include machine-learning researchers, data-systems builders, software reliability advocates and people who insisted that model performance must be evaluated in the real world.

The people below helped create the ideas and infrastructure that MLOps engineers use. Their contributions range from algorithms to operating practices, and their stories show why deploying a model is an engineering and social decision, not merely a training run.

The all-time podium

🥈
Jeff Dean
United States
2
🥇
Fei-Fei Li
United States
1
🥉
Daphne Koller
United States
3

“Data infrastructure can change what models are possible long before a new algorithm does.”

Fei-Fei Li

The eight who reached the top

1

Fei-Fei Li

United States · b. 1976

Li led the ImageNet project, whose large labeled dataset and benchmark transformed computer-vision research and exposed the importance of data infrastructure to machine learning.

The story

ImageNet began in 2006 as a project to organize millions of labeled images according to WordNet categories. Its 2012 competition became a turning point when a deep neural network dramatically improved results, accelerating the compute, data and deployment challenges later handled by MLOps teams.

“Data infrastructure can change what models are possible long before a new algorithm does.”

ImageNet launched
2006
Landmark contest
2012
Field
Computer vision
2

Jeff Dean

United States · b. 1968

Dean co-authored foundational Google systems papers and helped lead TensorFlow, linking large-scale distributed systems to practical machine-learning infrastructure.

The story

Google's 2016 release of TensorFlow as open source gave researchers and production engineers a common framework for describing and executing many machine-learning workloads. Its spread highlighted how frameworks, tooling and deployment environments influence which ideas can move from research into products.

“A shared platform turns isolated technical capability into a repeatable organizational capability.”

TensorFlow released
2015
Focus
Distributed systems
Institution
Google
3

Daphne Koller

United States · b. 1968

Koller made major contributions to probabilistic modeling and co-founded Coursera and Insitro, connecting machine learning research to large operational organizations.

The story

At Stanford, Koller co-authored the textbook Probabilistic Graphical Models, which gave students and practitioners a systematic vocabulary for uncertainty and structured prediction. Her later work building organizations around online learning and drug discovery illustrates the systems needed around models.

“A model's value depends on the organization that can use, test and sustain it.”

Textbook
Probabilistic Graphical Models
Co-founded
Coursera
Focus
ML applications
4

Cynthia Rudin

United States · b. 1976

Rudin's work on interpretable machine learning challenges teams to use transparent models in high-stakes settings where explanations are required.

The story

Rudin has argued that for high-stakes decisions, an opaque model paired with a post-hoc explanation is often a poor substitute for a model designed to be interpretable from the start. The argument directly affects MLOps evaluation and model-approval practices.

“When a decision affects people deeply, explainability is a design requirement, not a dashboard add-on.”

Focus
Interpretable ML
Use cases
High-stakes decisions
Role
Researcher
5

Martin Kleppmann

United Kingdom · b. 1987

Kleppmann's writing on distributed data systems gave engineers a practical framework for replication, consistency and stream processing—problems at the heart of reliable ML data pipelines.

The story

His 2017 book Designing Data-Intensive Applications translated hard-won distributed-systems lessons into a widely read engineering guide. MLOps engineers use the same questions about lineage, latency, correctness and failure when building training and inference data flows.

“Data pipelines are production systems; design them for disagreement, delay and failure.”

Book published
2017
Focus
Data systems
Audience
Production engineers
6

Rachel Thomas

United States · b. 1981

Thomas co-founded fast.ai and the nonprofit fastai community, emphasizing practical deep-learning education and the social consequences of deployed models.

The story

Through fast.ai courses, Thomas and collaborators made practical deep-learning instruction accessible to many learners outside elite research labs. Her later advocacy on accountability has reinforced the idea that deployment teams must measure harms, not only accuracy.

“Making a system easier to build does not remove the duty to understand who it affects.”

Co-founded
fast.ai
Focus
Practical ML
Theme
Accountability
7

Chip Huyen

Vietnam / United States · b. 1994

Huyen's writing and teaching on designing machine-learning systems gave the MLOps field a practical vocabulary for data distribution, deployment and monitoring.

The story

Her 2022 book Designing Machine Learning Systems collected operational lessons—data distribution shifts, feedback loops, evaluation and infrastructure—that had often been scattered across engineering teams. It became a common reference for people trying to make MLOps concrete.

“Treat deployment as the beginning of a model's evidence, not the end of its development.”

Book published
2022
Focus
ML systems
Contribution
MLOps education
8

Timnit Gebru

Ethiopia / United States · b. 1982

Gebru's research on dataset documentation and harms from large language models made evaluation, provenance and governance central concerns for production AI teams.

The story

Gebru co-authored the 2018 “Datasheets for Datasets” paper, proposing standardized documentation for datasets' motivation, composition and collection. The idea gives operational teams a way to ask what a model's training material represents before deploying it in a new setting.

“If data has no documented origin and limits, model confidence is not evidence.”

Datasheets paper
2018
Focus
AI accountability
Practice
Dataset documentation

Bars are scaled to the leader in this list.

Comparison Lab

Toggle names on and off — every bar rescales to the leader of your selection.

5 / 8

The argument

The boundary between MLOps, ML engineering, data engineering and platform engineering is contested because companies organize the same work differently. The useful question is not the title but who owns reproducibility, deployment, monitoring and model risk.

Teams also debate whether every model needs a complex platform. Small, low-risk applications may be safer with a simple, well-documented pipeline; overengineering can create its own operational debt.

Similar professions

Closest neighbours on the six-score profile — not the same field only.

Continue exploring

Keep exploring

More in Engineering & Technology