

Introduction to Machine Learning in Epidemiology with R
How can machine learning help us work with epidemiological data? And how do we know whether a predictive model is actually useful?
Join R-Ladies Rome for a practical, two-hour introduction to machine learning in epidemiology using R.
In this workshop, Federica Gazzelloni will introduce the main ideas behind supervised machine learning through an applied epidemiological example. Rather than focusing on a long list of algorithms, we will follow the complete machine-learning workflow: from defining an epidemiological question to training, evaluating and interpreting predictive models.
We will explore how different models approach the same prediction problem, starting with logistic regression as a baseline and moving to decision trees and random forests.
Using R, we will look at how to define a classification task, train models, generate predictions and evaluate their performance on unseen data.
Particular attention will be given to model evaluation and interpretation.
Throughout the workshop, we will also consider an important distinction for epidemiological research:
Prediction is not the same as inference, and predictive importance does not imply causation.
The workshop is inspired by the recent Machine Learning in Epidemiology study by Wright et al. (2026) and connects with Federica's book, Health Metrics and the Spread of Infectious Diseases: Machine Learning Applications and Spatial Modelling Analysis with R (CRC Press, 2025), where she introduces machine-learning applications in health and infectious-disease research using the mlr framework.
During the workshop, we will use the modern mlr3 ecosystem and discuss how machine-learning workflows in R have evolved from mlr to mlr3.
What we will cover
What machine learning means in an epidemiological context
From an epidemiological question to a prediction task
Preparing data for machine learning
Logistic regression as a baseline model
Decision trees and random forests
Training and evaluating models with
mlr3Cross-validation and performance on unseen data
Sensitivity, specificity, confusion matrices and ROC/AUC
Variable importance and model interpretation
Prediction versus explanation and causation
Limitations, bias and data quality in epidemiological machine learning
Who is this workshop for?
The workshop is designed for R users interested in epidemiology, public health, health data or machine learning. Basic familiarity with R and data analysis is useful, but no previous machine-learning experience is required.
The session will combine explanation, live R coding and discussion, with an emphasis on practical and reproducible analysis.
About the instructor
Federica Gazzelloni is an actuary, statistician, data scientist, author and instructor, and the founder and organiser of R-Ladies Rome. Her work spans health metrics, statistical modelling, machine learning, reproducible research and R.
She is the author of Health Metrics and the Spread of Infectious Diseases: Machine Learning Applications and Spatial Modelling Analysis with R, published by CRC Press in 2025.
R-Ladies
R-Ladies is a worldwide organisation promoting gender diversity in the R community. R-Ladies Rome provides a welcoming space to learn, share knowledge and connect with people interested in R, data science, statistics and reproducible research.
Everyone is welcome to attend, regardless of gender identity or level of experience.