Jev calibration

Confidence only helps when it matches observed outcomes.

Inspect reliability, expected calibration error, and confidence buckets on one recorded dataset before turning a Jev confidence value into automation.

Read confidence as evidence—not certainty.

Reliability curve

Compare predicted confidence with observed correctness across buckets and keep sample counts visible.

Task semantics

For Noul, values near zero can represent a confident negative—not weak confidence. The report preserves that distinction.

Small-sample guardrail

Insufficient rows or class coverage trigger EXPLORATORY / INSUFFICIENT_SAMPLE instead of a production recommendation.