Writing
On evaluation design, data quality, and the architecture around AI systems that have to actually work.
Precision 1.0, recall 1.0, and the artifact was still wrong
A pipeline scored a perfect parity match and failed the artifact anyway. Here is why, and why a confidence score should be arithmetic a person can re-run.
Read post-
The SLM scored 0.727, the LLM 0.90, and the comparison was noise
A fine-tuned small model against a large one on the same governance task. The headline gap was inside the error bars. The useful signal was which way each model got things wrong.
Read -
Measuring success of a model - practical case of anomaly detection on financial data
Why ROC AUC flatters rare-event detectors, and why the alert budget ended up mattering more than the model.
Read