Notes & findings

Writing

On evaluation design, data quality, and the architecture around AI systems that have to actually work.

Latest · 5 min read

Precision 1.0, recall 1.0, and the artifact was still wrong

A pipeline scored a perfect parity match and failed the artifact anyway. Here is why, and why a confidence score should be arithmetic a person can re-run.

evaluationqualityknowledge-base
Read post