THE PROBLEM PATTERN
AI programmes often start measuring after the wrong assumptions are already embedded.
Teams may compare models before understanding whether datasets represent the intended population, whether labels are consistent, whether sources are comparable or whether the selected metric reflects the costly failure.
Data quality, model quality and system quality are related. They are not interchangeable.
MY CONTRIBUTION
Turn research questions into architecture and acceptance gates.
- Developed profiling approaches for high-dimensional binary-feature datasets.
- Examined source differences and feature-frequency distributions before modelling.
- Created BenchMetrics methods and meta-metrics to test performance-measure robustness.
- Proposed systematic ML processes with dataset-quality comparison as a control gateway.
- Applied labelling, error analysis and feature-level evaluation in production mapping automation.
THE DECISION METHOD
Move from raw data to a defensible acceptance decision.
Establish fitness for purpose
Define the intended use, population, decision consequence and unacceptable failure.
Profile and compare sources
Expose distributions, sparsity, duplication, coverage, label quality and source-specific bias.
Separate quality layers
Distinguish data fitness, model behaviour and system-level performance.
Select metrics against failure modes
Justify measures against class distribution, baseline behaviour and decision risk.
Define acceptance evidence
Connect tests, thresholds, exceptions, human oversight and operational monitoring.
OUTCOME
A route from dataset uncertainty to an architecture decision.
The method reduces the risk of building a sophisticated model on poorly understood data or reporting success through a measure that does not reflect the real objective.
SELECTED EVIDENCE
Peer-reviewed methods and applied production practice.
- BenchMetrics and BenchMetrics Prob performance-instrument benchmarking.
- Q1 journal work on feature-space distributions and dataset quality.
- PToPI knowledge representation of performance measures.
- TasKar calculation and visualisation tool.
- Production feature-level evaluation and labelling-control experience.
Relevant capabilities: data profiling · AI readiness · evaluation design · metric selection · quality gates · acceptance evidence · responsible ML process.