#quality
2 posts
-
Precision 1.0, recall 1.0, and the artifact was still wrong
A pipeline scored a perfect parity match and failed the artifact anyway. Here is why, and why a confidence score should be arithmetic a person can re-run.
Read -
The SLM scored 0.727, the LLM 0.90, and the comparison was noise
A fine-tuned small model against a large one on the same governance task. The headline gap was inside the error bars. The useful signal was which way each model got things wrong.
Read