Description
Turn messy data into a defensible, analysis-ready data release
Real-world data cleaning involves more than correcting obvious errors. You must understand the source data, define clear quality rules, preserve evidence, document each decision, validate the results, and determine whether the cleaned data is safe to release.
In this hands-on CDI Practical, you will use Python to complete an end-to-end data-cleaning and validation project with a realistic synthetic service-delivery dataset. The dataset contains common quality problems involving missing values, duplicate records, inconsistent categories, invalid dates, malformed identifiers, and out-of-range numeric values.
By the end, you will have completed a reproducible data-quality workflow and created clear evidence of what changed, why it changed, which records were withheld, and whether the resulting dataset was ready for analysis.
What you will practice
- Profiling the source dataset and assessing its structure, fields, and data types
- Creating a source-row key and input-file fingerprint for traceability
- Normalizing column names safely and detecting naming collisions
- Assessing missing values and defining defensible treatment rules
- Detecting true duplicate records without losing source-row traceability
- Standardizing text, categorical, and yes-or-no values
- Validating dates while preserving the original source values
- Checking identifiers, numeric fields, and acceptable ranges
- Recording data-quality issues with rule, severity, and resolution status
- Separating analysis-ready records from records requiring further review
- Comparing data quality before and after cleaning
- Applying a formal publish-or-quarantine release gate
- Exporting reproducible outputs and a machine-readable run manifest
What you will produce
- A cleaned, analysis-ready CSV dataset
- A separate quarantined-records dataset
- A structured data-quality issue register
- A before-and-after data-quality scorecard
- Documented cleaning and validation decisions
- Automated validation results and final assertions
- A formal data-release decision
- A reproducibility manifest containing the input-file fingerprint
- Summary tables, quality-control files, and portfolio-ready evidence
What is included
- A browser-based, step-by-step CDI Practical
- A realistic synthetic source dataset
- Reusable Python cleaning and validation code
- Guided exercises, worked examples, and professional decision points
- Example quality-control outputs and documentation templates
- Access through the download link supplied after purchase
- The option to save a personal PDF copy using your browser’s Print → Save as PDF feature
Who this practical is for
- Data science and data analysis learners
- Monitoring and evaluation professionals
- Researchers working with tabular data
- Data-management and quality-assurance practitioners
- Professionals building evidence for a data portfolio
What you need
- A computer with a Python environment
- Jupyter Notebook, JupyterLab, VS Code, or another environment for running Python code
- Basic familiarity with Python and tabular data
- No advanced data-cleaning experience
Important information
The practical uses synthetic data created for learning. It contains no real participant or client records.
This is a digital product. No physical item will be shipped.




