Reliability diagrams: plot LLM confidence vs accuracy

Read a reliability diagram for an LLM classifier from Curva's reliability bins or dashboard, and spot overconfident, underconfident and leaning models.

A reliability diagram plots a classifier's stated confidence against its actual accuracy. Group the answers into bins by confidence, then put each bin's average confidence on the x axis and the share of its answers that were right on the y axis. A calibrated model sits on the diagonal. Points below the diagonal mean overconfident, points above mean underconfident. For an LLM classifier you can read the bins straight from Curva's calibration report, or open the dashboard, which draws the diagram with the diagonal for you. This post shows where the bins come from, how to read each shape, and how to turn the curve into a min_confidence cut-off.

Bins, the diagonal, and which side means overconfident

The idea is simple. If a model says 0.8 on a group of answers, about 80% of them should be right. Check that at every confidence level, and you have a calibration curve. Curva is named after that curve: Curva means "curve".

stated confidence observed accuracy 0 1 1 overconfident: below underconfident: above calibrated: on the line
Figure 1. A reliability diagram for an LLM classifier Made-up example: points below the diagonal are overconfident, points above are underconfident, and a calibrated model follows the dashed line.

How to read it:

  • **Below the diagonal:** the model claims more than it delivers. A bin at 0.9 that is right far less often is the classic LLM failure, and the most dangerous one, because high-confidence answers are the ones you automate.
  • **Above the diagonal:** the model is right more often than it says. Safe, but wasteful: you send answers to people that you could have automated.
  • **On the diagonal:** confidence means what it says.

The distance from the diagonal, weighted by how many answers sit in each bin, is the expected calibration error (ECE). The diagram shows you where that error lives.

Getting the bins: client.calibration and the reliability field

Every Curva answer needs a true label before it can appear in a reliability diagram. Send labels with feedback when you learn them, then ask for the report:

python
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])

Over HTTP it is GET /v1/calibration?question=<key>&project=<name>. The report covers the most recent wording of the question. Both before and after hold accuracy, ece, brier, automated and accuracy_when_automated, plus reliability bins: the predicted confidence against the observed accuracy, which is what you plot. Feed the reliability bins of either block to any plotting tool you like, or skip the plotting entirely and use the dashboard.

The dashboard's calibration panel draws the reliability diagram

curva serve has a built-in page at /dashboard. Its Calibration view reads GET /v1/calibration for one question and shows the labels collected against the 30 needed, the fitted calibrator, accuracy, ECE, Brier and automation before and after calibration, and a reliability diagram: predicted confidence against observed accuracy, with the diagonal a calibrated model follows.

The page itself holds no data, so it is served without a key. Every data request it makes goes to the API and needs a key like any other client. Give the dashboard its own key with curva keys create --name dashboard, so you can revoke it without touching production clients. On anything but localhost, put the server behind TLS.

Before vs after: why the after curve is held out

The report draws two curves. before is the model's raw probabilities. after is what calibration would do, measured honestly: each half of the labels is calibrated by a fit on the other half. A curve fitted and tested on the same labels would always look good. A held-out curve can look worse, and when it does, Curva doesn't apply the calibrator.

Curva applies a calibrator only when it clearly beats the raw probabilities on held-out labels, in both log-loss and Brier score. If the after curve barely differs from before, expect answers to stay raw with calibrated: false. That is the right outcome for a model that already sits near the diagonal.

For a sense of scale on public data: in Curva's runs of 2026-10-01, Gemini flash-lite on AITA had a raw ECE of 0.404 and 0.149 after calibration, held out (n = 100). On BoolQ the same model's ECE stayed at 0.081 before and after (n = 127), because calibration couldn't show a clear gain.

Few answers per bin: what 30 labels can and cannot show

A reliability diagram needs answers in each bin to mean anything. Thirty labels is Curva's minimum to fit a calibrator, but 30 labels spread over several bins leaves only a handful per bin. A bin with four answers can sit far from the diagonal by chance.

So read an early diagram for its shape, not its points. A curve that sits below the diagonal across every bin says "overconfident" even with few labels. A single odd bin says little. Curva handles this on its side too: with 30 labels a calibrator moves probabilities only about a third of the way to what the labels suggest, and nearly all the way at 1,000.

Two habits make the diagram trustworthy faster. Label confident answers as well as the ones people reviewed, so the high bins fill up. And keep the question's wording fixed, because a reworded question starts a fresh calibration with no labels.

From the curve to a min_confidence cut-off

The point of the diagram is a decision: where do you stop automating? Find the lowest confidence where the curve meets the accuracy you need, and set min_confidence there. Answers below it come back with abstain: true and go to a person.

The report gives the shortcut for the most common cut-off. accuracy_when_automated is the accuracy on answers at or above 0.9, and automated is the share of answers that reach it. If that accuracy is high enough for you, 0.9 is a usable threshold. If it isn't, the diagram shows how far up you would have to move it, or whether the model needs calibration first.

Before calibration, the threshold is a guess about the model. Once the question is calibrated, the curve is close to the diagonal and the threshold means what it says.

Next steps

The docs explain the curve in probabilities and calibration and the dashboard guide. For the number the diagram summarises, read expected calibration error. For the mechanism, see LLM calibration explained, and for what to automate at each threshold, selective accuracy for LLMs. The full picture is in LLM classification confidence scores you can act on.