Reliability curve
Compare predicted confidence with observed correctness across buckets and keep sample counts visible.
Inspect reliability, expected calibration error, and confidence buckets on one recorded dataset before turning a Jev confidence value into automation.
Compare predicted confidence with observed correctness across buckets and keep sample counts visible.
For Noul, values near zero can represent a confident negative—not weak confidence. The report preserves that distinction.
Insufficient rows or class coverage trigger EXPLORATORY / INSUFFICIENT_SAMPLE instead of a production recommendation.