All topics

#Evaluation

Metrics whose assumptions match the problem, splits that don’t flatter the model, and trajectory-level evaluation for agents — numbers a person can actually act on.

1 post