Introduction: Machine learning (ML) models have been proposed for microbiological risk assessment (MRA) of foods to integrate complex data. However, external validation and operational readiness remain limited. This review critically synthesizes the application of supervised ML to the MRA of animal-derived foods, with an emphasis on validation strategies.
Methods: Following PRISMA 2020, we searched Scopus, Web of Science, and PubMed for original studies. After duplicate removal, articles were screened using Rayyan, followed by full-text review. Inclusion criteria: (i) supervised ML for prediction; (ii) microbiological outcome (pathogen presence/concentration, antimicrobial resistance); (iii) animal-derived food matrices or production environments; (iv) performance metrics and validation reported. Including 21 studies, total.
Results: Random Forest ranked best in 9/21 studies, boosting models (XGBoost, LightGBM, GBM) in 6/21, and neural networks in 3/21. Accuracy (12/21) and AUC-ROC (9/21) were the most reported. Only 7 studies reported F1-score, and 2 MCC – metrics are more suitable for imbalanced data. K-fold cross-validation was common (17/21), yet true external validation (temporal/geographic) occurred in only 8 studies, and only Mi et al. (2025) performed prospective validation on unseen data. Data leakage was discussed in 3 articles, but potentially affected 12 investigations that split data by record without grouping by animal or farm. Meteorological variables were the dominant predictors (4/21), while genomic data (2/21) enabled strain-level resolution. Most studies (15/21) addressed Exposure Assessment; only Njage et al. (2020) integrated ML into a full quantitative MRA, showing that ignoring genetic heterogeneity can lead to over‑ or underestimation of risk.
Conclusion: Despite high predictive performance (AUC > 0.85, R² > 0.93), gaps persist: external validation is missing in 62% of studies, data leakage risk in 57%, and code is unavailable in 76%. Recent advances (2023–2025) in explainable AI, rigorous temporal validation, and operational tools show progress, but stricter methodological standards are needed for regulatory uptake.