Skip to content

Calibration

A model is calibrated when its 0.8 answers are right about 80% of the time. Most models are close but not exact, and the gap depends on the domain and the language. JevOss fits two fixes on held-out data you trust.

jevoss --endpoint http://127.0.0.1:8000 calibrate my_labelled.jsonl --out calibration.json

Temperature per question type and language

One temperature T rescales every distribution: softmax(log p / T). T > 1 softens an overconfident model, T < 1 sharpens a timid one. The top option never changes.

JevOss fits T by minimising negative log-likelihood, separately for each question type and language when a group has at least 150 decisions, and falls back to the type, then to one global value. The output uses the same keys the JevAlt engine reads (choice:tr, score, default).

Conformal sets

Split conformal prediction turns probabilities into an option set with a coverage guarantee: with coverage = 0.9, the set contains the correct answer for at least 90% of new decisions drawn from the same distribution as the calibration data. Clear cases get a single option, unclear cases get two or three.

JevOss stores one threshold per group and coverage level. A JevAlt server returns the set when the request asks for it:

{"state": "...", "questions": {"team": {...}}, "coverage": 0.9}
{"team": {"choice": "billing", "probabilities": {"billing": 0.71, "technical": 0.24, "sales": 0.05}, "set": ["billing", "technical"]}}

The guarantee holds on average over new data from the same distribution. Recalibrate when your traffic changes.

Reading the numbers

Metric Good direction What it tells you
accuracy higher top option matches the label
Brier lower squared error of the whole distribution
NLL lower how surprised the model is by the truth
ECE lower gap between stated confidence and hit rate (15 equal-mass bins)
KL to gold lower distance to a reference distribution, when the suite has one
AUROC higher whether confidence separates right from wrong answers
coverage@5% higher share of decisions you can automate while keeping errors under 5%