Jev evaluation

Evaluate Jev on the labels and edge cases you actually own.

A public example cannot tell you how one Jev version handles your classes, questions, and failure costs. Build a dated decision record from your own labeled rows.

An evaluation is more than one accuracy number.

Task-correct metrics

Choice and Noul use different metrics. The report does not force every task into one generic score.

Row-level evidence

Review errors, fallbacks, and class behavior instead of accepting an aggregate without its failure cases.

Versioned context

Record the model route, question hash, dataset hash, threshold method, and measured limitations.