Description
Introduction to Machine Learning Practical is a hands-on guide to building a complete, reproducible, and responsible binary-classification workflow with Python and scikit-learn.
You will work with cleaned synthetic service-delivery data and progress from defining the prediction task to comparing models, evaluating an untouched holdout set, selecting a decision threshold, checking subgroup performance, and reporting results responsibly. The workflow compares a prevalence-only baseline, logistic regression, and a shallow decision tree using mixed-type preprocessing pipelines and cross-validation.
This practical goes beyond reporting accuracy. You will examine precision, recall, F1, ROC AUC, average precision, confusion matrices, predicted probabilities, and threshold trade-offs while keeping prediction distinct from explanation and causal inference.
What You Will Learn
- Distinguish prediction from explanation and causal inference.
- Define the target, features, observation unit, and prediction moment.
- Identify identifier, proxy, temporal, and preprocessing leakage.
- Create reproducible stratified training and holdout sets.
- Build mixed-type preprocessing with scikit-learn pipelines.
- Compare models using cross-validation and an untouched holdout set.
- Interpret accuracy, precision, recall, F1, ROC AUC, average precision, confusion matrices, predicted probabilities, and decision thresholds.
- Check subgroup performance and communicate responsible-use boundaries.
What You Will Produce
- A prediction-task, feature, and leakage review.
- Cross-validation and holdout performance tables.
- Confusion-matrix, ROC, precision–recall, target-balance, and threshold figures.
- Subgroup performance, coefficient, threshold, and prediction tables.
- A saved end-to-end model pipeline.
- A model card, model-decision log, and portfolio-ready project statement.
Included
- Online step-by-step CDI Practical
- Downloadable project files
- Deterministic synthetic dataset generator
- Complete standalone Python workflow
- Reference tables, figures, and saved example pipeline
- Learner documentation and reproducible packaging scripts
Who This Practical Is For
This practical is designed for students, researchers, analysts, professionals, mentors, and learners seeking a clear introduction to reproducible and responsible machine learning.
Tools: Python, pandas, NumPy, scikit-learn, Matplotlib, seaborn, joblib, and Quarto.
Important: This practical is intended for education, portfolio evidence, and local technical testing. Its synthetic data and demonstration outputs are not validated for operational, clinical, or other real-world decisions affecting people.




