STATE OF PRODUCTION ML · 2024

The State of Production ML in 2024

What 177 practitioners reported from production: CI/CD supported 69% of deployment workflows, AWS served more than half of cloud users, monitoring was the top challenge at 45%, and PyTorch led the library field at 36%.

TOP CHALLENGE0145%

selected monitoring and observability as a production challenge

LEADING LIBRARY0236%

primarily used PyTorch, Lightning or Fast.ai

DEPLOYMENT FOUNDATION0369%

reported CI/CD support for machine learning deployment

SURVEY RESPONSES04177

submitted rows, with question bases varying by nonblank answers

01 / ML CONTEXT

ML context

What teams built, where they ran it and what slowed the route from experiment to production.
FINDING 01

PyTorch led a divided library field.

PyTorch, Lightning and Fast.ai accounted for 36% of answers. Scikit-learn followed at 31%.

  1. PyTorch/Lightning/Fast.ai36%
  2. scikit-learn31%
  3. XGBoost10%
  4. TensorFlow8%
  5. CatBoost5%
  6. LightGBM4%
  7. LangChain1%
  8. Dspy1%
View all 16 options
  1. gpflow1%
  2. Nixtla1%
  3. Nixtla, darts, sktime, gluonts1%
  4. Non1%
  5. Perpetual1%
  6. Pymc1%
  7. Stable-Baselines31%
  8. Triton1%
FINDING 02

LLMs became the leading reported modality.

LLM work reached 47%, with time series at 42% and tabular machine learning at 39%.

Multiple selections allowed
  1. LLMs46%
  2. Time Series / Forecasting42%
  3. Tabular39%
  4. Text / NLP (Non-LLM)38%
  5. Recommender Systems32%
  6. Image / Computer Vision25%
  7. Search17%
  8. Causal Inference16%
View all 13 options
  1. 3D data (volumetric, mesh...)1%
  2. Graphs1%
  3. Regression1%
  4. Reinforcement Learning1%
  5. RL1%
FINDING 03

Recommenders led a broad use-case mix.

Recommender systems were selected by 39% of respondents. Demand forecasting followed at 32%.

Multiple selections allowed
  1. Recommender systems39%
  2. Demand Forecasting32%
  3. Fraud28%
  4. Risk27%
  5. Search26%
  6. Marketing Intelligence23%
  7. Pricing21%
  8. Classification1%
View all 60 options
  1. Digitalisation1%
  2. Optimization1%
  3. .1%
  4. Agentic systems1%
  5. Agents1%
  6. agents assists1%
  7. Anomaly detection1%
  8. Anomaly Détection1%
  9. Augmentation1%
  10. Autonomous driving1%
  11. Av1%
  12. biotech use cases1%
  13. Chatbot1%
  14. Climate Modelling1%
  15. Computer Vision1%
  16. content generation1%
  17. Cost1%
  18. Customer Satisfaction1%
  19. Data analysis and enhancement1%
  20. Decision support1%
  21. defect detection1%
  22. Defect detection1%
  23. document (pii) anonymisation1%
  24. Document entity extraction1%
  25. Drug Discovery1%
  26. Education1%
  27. finance1%
  28. Generating synthetic data1%
  29. Geospatial AI1%
  30. Healthcare1%
  31. ICR/OCR1%
  32. Image Classification1%
  33. Information extraction1%
  34. Insurance1%
  35. Meteorology1%
  36. OCR based tasks1%
  37. Pharmaceuticals1%
  38. Predictive maintenance1%
  39. process optimization1%
  40. project duration forecasting1%
  41. Property Prediction (Molecular)1%
  42. quality assurance1%
  43. robotics1%
  44. speech recognition1%
  45. stock1%
  46. Strategy1%
  47. Support automation1%
  48. Sustainability1%
  49. Synthetic data generation1%
  50. Text Summarization1%
  51. translation and multilinguality1%
  52. Visual Detection in Video1%
FINDING 04

Most production cycles stayed below six months.

Thirty-one percent reported one to three months, while 26% took three to six months and 24% less than a month.

  1. Less than 3 months31%
  2. Less than 6 months26%
  3. Less than a month24%
  4. Less than a week7%
  5. More than a year6%
  6. Less than a year5%
FINDING 05

AWS held a clear cloud lead.

Amazon Web Services accounted for 54% of answers. Azure followed at 21% and Google Cloud Platform at 16%.

  1. Amazon Web Services54%
  2. Azure21%
  3. Google Cloud Platform16%
  4. On-prem4%
  5. Databricks1%
  6. AWS & GCP1%
  7. Digital Ocean1%
  8. GCP and AWS equally1%
View all 10 options
  1. IBM1%
  2. No cloud platform used1%
FINDING 06

Monitoring was the defining production challenge.

Monitoring and observability was selected by 45%. Tooling gaps and access to training data followed at roughly one third.

Multiple selections allowed
  1. Machine learning monitoring and observability45%
  2. Gaps in tooling and support for model productionisation33%
  3. Access to relevant data for training32%
  4. Building production-grade machine learning and data pipelines32%
  5. Inconsistency of training and experimentation environments31%
  6. Showcasing business impact and business value29%
  7. Lack of specialised engineers27%
  8. Governance and domain risks17%
View all 17 options
  1. Lack of specialised data scientists11%
  2. Machine learning security7%
  3. Big organisational migrations (of Datasets, MLP and Data platforms and mroe)1%
  4. client requirements1%
  5. Inconsistency of training and production environments1%
  6. lack of vision (what happens after productionising)1%
  7. Model consistency (re-training / catastrophic forgetting)1%
  8. reservedness from stakeholders1%
  9. Unit economics1%

02 / PLATFORMS & TOOLS

Platforms & tools

The systems practitioners chose across tracking, data, training, serving, monitoring and foundation models.
FINDING 07

MLflow set the pace for experiment tracking.

MLflow represented 48% of respondents using a tracking or registry tool. In-house systems followed at 16%.

  1. MLflow48%
  2. Custom Built In-house tool16%
  3. Weights & Biases12%
  4. Spreadsheets8%
  5. ClearML4%
  6. Data Version Control (DVC)3%
  7. Comet1%
  8. Azure ML Service, Sagemaker1%
View all 17 options
  1. Domino1%
  2. GCS and Clearml1%
  3. Google MetaDataStore1%
  4. Hopsworks1%
  5. omega-ml1%
  6. Optuna1%
  7. Sagemaker1%
  8. Snowflake1%
  9. Vertex AI Experiments1%
FINDING 08

Feature stores were predominantly built in-house.

In-house systems accounted for 62% of feature-store answers. FEAST was the leading packaged option at 10%.

  1. Custom Built In-house tool62%
  2. FEAST10%
  3. Databricks7%
  4. Hopsworks7%
  5. None3%
  6. DynamoDB2%
  7. Snowflake2%
  8. Tecton2%
View all 12 options
  1. Dataiku1%
  2. Fennel1%
  3. none1%
  4. selft built / Databricks1%
FINDING 09

Vector database choices remained fragmented.

Pinecone led at 14%, narrowly ahead of Milvus at 12%. No option held a dominant share.

  1. Pinecone14%
  2. Milvus12%
  3. Custom Built In-house tool10%
  4. Weaviate7%
  5. Databricks5%
  6. Elasticsearch5%
  7. pgvector5%
  8. Azure AI Search4%
View all 32 options
  1. None4%
  2. OpenSearch4%
  3. FAISS2%
  4. PostgreSQL2%
  5. Qdrant2%
  6. ActiveLoop1%
  7. Azure Cognitive Search1%
  8. azure's offering in ai search1%
  9. Chroma1%
  10. ChromaDB1%
  11. DynamoDB1%
  12. Elastic Search1%
  13. Elasticsearch / Opensearch1%
  14. gcp1%
  15. Hopsworks1%
  16. LanceDB1%
  17. LanceDB / Turbopuffer1%
  18. MongoDB1%
  19. NA1%
  20. nil1%
  21. No vector database1%
  22. opensearch / Azure AI Search1%
  23. Opensearch AWS1%
  24. opensearch, elasticsearch1%
FINDING 10

Airflow was the default workflow orchestrator.

Airflow accounted for 41% of answers. In-house systems followed at 17% and Argo Workflows at 11%.

  1. Airflow41%
  2. Custom Built In-house tool17%
  3. Argo Workflows11%
  4. Prefect5%
  5. Databricks4%
  6. Amazon Glue2%
  7. Databricks Workflows2%
  8. Celery1%
View all 30 options
  1. Kubeflow1%
  2. Metaflow1%
  3. Abinitio1%
  4. aws step function1%
  5. AWS Step Functions1%
  6. Azure data factory, Databricks workflows, azureML pipelines1%
  7. Azure devops1%
  8. Azure ML Pipelines1%
  9. ClearML1%
  10. Dagster1%
  11. Databricks workflow1%
  12. Dataiku1%
  13. Dbt1%
  14. GitHub Actions1%
  15. kubeflow1%
  16. kubeflow and ray1%
  17. kubeflow pipelines1%
  18. Meta flow1%
  19. NiFi1%
  20. None1%
  21. snakemake1%
  22. Snowflake1%
FINDING 11

In-house platforms led model training.

In-house systems represented 34%, followed by Databricks at 23% and Amazon SageMaker at 17%.

  1. Custom Built In-house tool34%
  2. Databricks23%
  3. Amazon SageMaker17%
  4. Google Cloud Vertex AI10%
  5. Azure ML Studio8%
  6. Domino1%
  7. Anyscale and also Slurm1%
  8. Argo Workflows1%
View all 17 options
  1. ClearML1%
  2. Hopsworks1%
  3. kubeflow1%
  4. Kubernetes, docker1%
  5. Local1%
  6. Metaflow1%
  7. Omega-ml1%
  8. Snowflake1%
  9. Weights and biases1%
FINDING 12

Python web frameworks dominated real-time serving.

FastAPI and Flask wrappers accounted for 47%. In-house systems followed at 16% and SageMaker at 12%.

  1. FastAPI/Flask Wrapper47%
  2. Custom Built In-house tool16%
  3. Sagemaker12%
  4. Databricks7%
  5. KServe3%
  6. Seldon Core2%
  7. BentoML1%
  8. Nvidia Triton Inference Server1%
View all 21 options
  1. Anyscale/Ray1%
  2. AWS Lambda1%
  3. Azure Managed Online Endpoints1%
  4. Azure ML1%
  5. Kubernetes1%
  6. LMDeploy1%
  7. Omega-ml1%
  8. SkyPilot1%
  9. Spring Boot1%
  10. tensorrt1%
  11. TorchServe1%
  12. Triton1%
  13. Triton Inference Server1%
FINDING 13

Most monitoring stacks were built internally.

In-house systems represented 52% of answers. Evidently was the leading packaged tool at 19%.

  1. Custom Built In-house tool52%
  2. Evidently AI19%
  3. Arize AI6%
  4. Neptune AI6%
  5. NannyML3%
  6. Azure ML2%
  7. Comet1%
  8. Domino1%
View all 17 options
  1. Fiddler AI1%
  2. Grafana1%
  3. Hopsworks1%
  4. mlcore1%
  5. Model metrics monitoring on Metabase1%
  6. Omega-ml1%
  7. Perpetual ML Suite1%
  8. Sagemaker1%
  9. WhyLabs1%
FINDING 14

Data platforms had no runaway winner.

Delta Lake led at 22%, followed by AWS Lake Formation at 18% and Snowflake at 17%.

  1. Deltalake22%
  2. AWS / Lakeformation18%
  3. Snowflake17%
  4. Custom Built In-house tool16%
  5. GCP / BigLake15%
  6. Azure DataLake8%
  7. Dremio1%
  8. IBM1%
View all 11 options
  1. MSSQL1%
  2. TDengine1%
  3. Trino1%
FINDING 15

OpenAI led managed foundation model services.

OpenAI accounted for 40% of answers. Azure AI followed at 21% and Amazon Bedrock at 13%.

  1. OpenAI40%
  2. Azure AI21%
  3. Amazon Bedrock13%
  4. Custom Built In-house tool9%
  5. Google Gemini6%
  6. Anthropic5%
  7. Ensemble1%
  8. fireworks1%
View all 12 options
  1. Mistral1%
  2. None1%
  3. Omega-ml1%
  4. Snowflake1%

03 / ORGANISATION & OPERATIONS

Organisation & operations

How production ML was organised, scaled and deployed across teams and infrastructure.
FINDING 16

Technology and finance anchored the sample.

Technology represented 21% of respondents and financial services 17%, with the remainder spread across many sectors.

  1. Technology21%
  2. Financial services17%
  3. Retail10%
  4. Healthcare8%
  5. Media & Entertainment7%
  6. Insurance5%
  7. Energy4%
  8. Transportation & warehousing4%
View all 34 options
  1. Telecommunications4%
  2. Food3%
  3. Automotive2%
  4. Consultancy2%
  5. Consulting1%
  6. Utilities1%
  7. Aerospace & defense1%
  8. Agriculture1%
  9. Biotech1%
  10. consulting1%
  11. Data Consulting1%
  12. Electronics1%
  13. Enviroment1%
  14. Fashion1%
  15. Food Delivery1%
  16. Government1%
  17. Hospitality1%
  18. Human Resources1%
  19. Manufacturing1%
  20. Market Research1%
  21. Marketplaces1%
  22. Pharma1%
  23. Pharmaceuticals1%
  24. Public Sector1%
  25. Saas1%
  26. Travel1%
FINDING 17

The sample spanned organisations of every scale.

Organisations with 50 to 250 employees were the largest group at 20%, followed by 250 to 1,000 at 18%.

  1. 50-250 employees20%
  2. 250-1,000 employees18%
  3. 5,000-50,000 employees17%
  4. 1,000-5,000 employees17%
  5. 10-50 employees10%
  6. 50,000+ employees9%
  7. Less than 10 employees8%
FINDING 18

Production fleets were usually measured in tens.

Two to five models was the largest group at 24%. Ten to twenty and twenty-one to one hundred followed at 20% and 19%.

  1. 2-524%
  2. 10-2020%
  3. 21-10019%
  4. 100-100014%
  5. 5-912%
  6. 1000+6%
  7. 14%
  8. 02%
FINDING 19

Planned fleets pointed toward continued expansion.

Twenty-five percent planned for 21 to 100 models. Five to nine followed at 20%, with ten to twenty at 18%.

  1. 21-10025%
  2. 5-920%
  3. 10-2018%
  4. 100-100014%
  5. 1000+12%
  6. 2-58%
  7. 13%
  8. 01%
FINDING 20

Central ML and data platforms were the operating core.

Central ML teams and data platform organisations were each reported by about 65% of respondents.

Multiple selections allowed
  1. Central Machine Learning Platform / Team65%
  2. Data Platform / Data Engineering Organisation65%
  3. A Developer Productivity Team which also covers machine learning26%
  4. AI Risk & Governance Function20%
  5. AI Inventory (Keeping track of all AI usecases and models)17%
FINDING 21

Batch remained central to inference workloads.

Thirty-five percent said fewer than one in ten models ran real-time inference. Twenty-two percent reported more than nine in ten.

  1. Less than 10%35%
  2. Between 10% and 30%24%
  3. More than 90%22%
  4. Between 50% and 90%20%
FINDING 22

CI/CD underpinned most deployment workflows.

Sixty-nine percent reported CI/CD support, while 58% had separate development, staging and production environments.

Multiple selections allowed
  1. CI/CD for continuous deployment69%
  2. Development-Staging-Production Environments58%
  3. A/B Tests for Models42%
  4. Progressive Rollouts30%
  5. Canary Deployments26%
  6. None1%
FINDING 23

ML engineers formed the largest practitioner group.

Machine learning engineers accounted for 44%, followed by data scientists at 23% and MLOps engineers at 19%.

  1. Machine Learning Engineer44%
  2. Data Scientist23%
  3. MLOps Engineer19%
  4. Software Engineer6%
  5. Business / Domain Practitioner3%
  6. Product Manager3%
  7. Data Engineer2%
  8. Data Analyst1%

04 / RESPONDENT PROFILE

Respondent profile

The role levels, ages, locations and identities represented in this survey.
FINDING 24

Individual contributors dominated the sample.

Junior to senior individual contributors represented 38%, with staff-level contributors at 31% and managers at 18%.

  1. Individual Contributor (Junior to Senior)38%
  2. Individual Contributor (Staff+)31%
  3. Manager18%
  4. Director/VP11%
  5. C-Suite2%
FINDING 25

Respondents clustered in their early thirties.

Ages 30 to 34 represented 35% of answers. Ages 35 to 39 followed at 22%.

  1. 30-3435%
  2. 35-3922%
  3. 25-2914%
  4. 40-4410%
  5. 45-498%
  6. 50-543%
  7. Prefer not to share3%
  8. 22-242%
View all 10 options
  1. 55-591%
  2. 18-211%
FINDING 26

Responses came from a distributed international community.

The USA was the largest consistently labelled group at 18%, followed by the Netherlands at 9% and the United Kingdom at 8%.

  1. USA18%
  2. Netherlands9%
  3. United Kingdom8%
  4. Germany8%
  5. India8%
  6. France5%
  7. Spain5%
  8. Japan3%
View all 40 options
  1. Switzerland3%
  2. Brazil2%
  3. Canada2%
  4. Israel2%
  5. Turkey2%
  6. Colombia2%
  7. india2%
  8. Portugal2%
  9. United States2%
  10. Australia1%
  11. Austria1%
  12. Bangladesh1%
  13. Czech Republic1%
  14. Denmark1%
  15. Estonia1%
  16. Finland1%
  17. FRANCE1%
  18. germany1%
  19. Italy1%
  20. Luxembourg1%
  21. netherlands1%
  22. nl1%
  23. Northern Ireland1%
  24. Norway1%
  25. portugal1%
  26. Romania1%
  27. Thailand1%
  28. The Netherlands1%
  29. UAE1%
  30. UK1%
  31. US1%
  32. Vietnam1%
FINDING 27

The gender imbalance was pronounced.

Ninety-three percent self-identified as male and 7% as female, with one non-binary response.

  1. Male93%
  2. Female7%
  3. Non-binary1%

05 / METHODOLOGY

Methodology

This baseline edition reports the 2024 answers without a prior-year comparison. Question-level response bases are summarised here so every percentage can be read in context.
RESPONSES177

Submitted rows in the 2024 source. Per-question nonblank respondent bases range from 81 to 170.

COLLECTION2024

Responses were collected through the State of Production ML survey. The report aggregates the supplied CSV at build time and publishes no raw-response table.

2025 EDITION AVAILABLE