Conformal prediction sets for LLM classification

Set coverage=0.95 and each answer carries a set of labels holding the right one at least 95% of the time. How split conformal works and what it needs.

Conformal prediction for an LLM classifier turns each answer into a short set of labels that contains the right one at a rate you choose, for example at least 95% of the time. It works on top of any model, however well or badly calibrated, because the threshold is set from your own labeled answers. In Curva you add coverage=0.95 to a question, send feedback as usual, and once the question has 30 labels every answer carries a set with that guarantee. This post shows how to read the set, builds the threshold by hand, and lists what the guarantee needs.

A promise instead of an average: coverage=0.95

A calibrated probability tells you how often answers like this one are right, on average. That is useful, but it is a statement about averages. Sometimes you need a promise instead: "the right answer is in this short list at least 95% of the time."

Curva implements split conformal prediction per question. You set coverage on a Choice, Multi or Noul question:

python
from curva import Curva, Choice

curva = Curva()
d = curva.decide(ticket, {"team": Choice("Which team?", ["billing", "technical", "sales"], coverage=0.95)},
                 project="support")

d["team"].choice       # "billing": the usual top answer, probabilities and confidence are all still there
d["team"].guaranteed   # True once the question has 30 labels
d["team"].set          # ["billing"], or ["billing", "technical"] when the model is torn

Everything else about the answer stays: the top choice, its confidence, the probability for every option. coverage only adds the set. In TypeScript the same option is { coverage: 0.95 } on the choice builder, and the set is d.answers.team.set.

Reading a conformal prediction set: one option, several, every option

SetMeaningWhat to do
One optionThe model is sure enough to meet the guarantee aloneAutomate
Several optionsThe truth is one of these, at the promised rateShow the options, or route to a person
Every optionThe model can't narrow it down for this inputRoute to a person

Several options is still useful. A support agent who sees "billing or technical" starts ahead of one who sees an unsorted queue. A review tool can show the two likely categories side by side.

A Noul's set is a subset of ["true", "false"], so a set with both values means the model can't tell.

Nonconformity 1 minus p(true) and q hat at rank ceil((n+1)c)

Curva builds the set with split conformal prediction, using the question's own feedback labels as the calibration set. Three steps:

  1. **Score every labeled answer.** For each past decision with a label, the nonconformity score is 1 − p(true answer): how little probability the model gave to what turned out to be right. A confident correct answer scores near 0. A confident wrong one scores near 1.
  2. **Find the threshold.** With n scores and coverage c, the threshold q̂ is the ⌈(n+1)·c⌉-th smallest score.
  3. **Build the set.** For a new answer, the set is every option with p ≥ 1 − q̂.

The probabilities are the ones /v1/decide returns: calibrated once the question has a calibrator, raw before. Two edge cases are handled. When the rank is past n (few labels and a high c), q̂ is 1 and the set is every option: the guarantee holds, it just isn't useful yet. The set is only empty when q̂ is 0, and then Curva returns the top answer, so there is always something to act on.

flowchart LR
  L["labeled answers"] --> S["score each: 1 minus p(true)"]
  S --> Q["q hat: the ceil((n+1)c)-th smallest score"]
  N["new answer probabilities"] --> K["keep options with p at least 1 minus q hat"]
  Q --> K
  K --> R["the set"]
Figure 1. How a conformal prediction set is built Labeled answers set the threshold; each new answer keeps every option at or above 1 minus q hat.

Worked example: 39 labels at coverage 0.9

Here is the arithmetic with made-up numbers. A question has 39 labels and you ask for coverage=0.9. The rank is ⌈40 × 0.9⌉ = 36, so q̂ is the 36th smallest of the 39 scores. Say that score is 0.55. Then the set is every option with probability at least 0.45. A new answer with billing 0.82, technical 0.12 and sales 0.06 gets the set ["billing"]. One with billing 0.48 and technical 0.47 gets ["billing", "technical"]. These numbers are invented to show the method.

Notice what happened in the second case. A plain top-answer rule would route it to billing, at a confidence below one half. The set says honestly that it is one of two.

A worse model gets bigger sets, not a broken promise

The guarantee doesn't depend on how good the model is. It needs only one thing: new inputs that look like the labeled ones, statistically. If your labels are a fair sample of your traffic, a new answer's score is equally likely to fall anywhere among the old scores. So the chance that it lands above the ⌈(n+1)·c⌉-th smallest is at most 1 − c. When it doesn't, the true answer is in the set.

The consequence is the line the Curva docs use: a worse model gets bigger sets, not a broken promise. A weak model spreads its probability around, its scores are high, q̂ is high, and sets grow. A strong model gets small sets. Coverage stays at the level you asked for either way. Calibration usually makes the sets smaller, because more honest probabilities rank options better, but the guarantee doesn't rely on it.

Two caveats keep this honest:

  • **It is a guarantee on average over your traffic**, not for each input. At 0.95 coverage, about one set in twenty will miss, and you won't know which.
  • **It needs representative labels.** If you only label the hard cases a person reviewed, the labels don't look like your traffic and the guarantee is off. Label a random sample too. If your traffic changes, the old labels stop representing it. The weekly drift report is the early warning.

Needs 30 labels and the same wording; Multi not guaranteed yet

What the guarantee needs:

  • **30 labels for this exact question in this project**, the same minimum as calibration. Before that, the answer has guaranteed: false and no set. It is still a normal answer, not an error.
  • **The same wording.** Labels, calibrator and guarantee belong to the exact question. Rewording starts over. Changing only coverage does not.
  • **Choice, Multi or Noul.** Other types, or a value outside (0, 1), get 422.
  • **Multi is not guaranteed yet.** Feedback labels a Multi as a whole, not option by option, so a Multi with coverage returns guaranteed: false for now.

Labels go in the usual way, with curva.feedback(d.id, "team", "technical").

One cost detail is worth knowing. coverage is not part of the question's fingerprint or of the decision cache key, because the set is worked out from stored labels after the model has answered. Asking the same question at 0.8 and at 0.95 costs one model call and returns two sets: a tight one for display, a safer one for automation.

Coverage and abstain together

coverage and min_confidence answer different questions, and they combine well:

  • abstain looks at the top answer: is its probability above min_confidence?
  • set looks at all the options: which must you keep to be right at the promised rate?

A common routing rule is to automate when the set has one option and abstain is false, and send everything else to review:

python
q = Choice("Which team?", ["billing", "technical", "sales"], min_confidence=0.8, coverage=0.95)
d = curva.decide(ticket, {"team": q}, project="support")
a = d["team"]
if a.guaranteed and len(a.set) == 1 and not a.abstain:
    route(ticket, a.choice)
else:
    review(ticket, candidates=a.set or [a.choice])

The review queue then feeds labels back, which builds the calibration set the guarantee rests on.

Next steps

The docs page is guaranteed accuracy. Next, read what the guarantee assumes in conformal prediction assumptions, how to pick a level in conformal coverage levels, and when to use a threshold instead in abstain vs conformal prediction. For how sets fit with calibration and abstain, see LLM classification confidence scores you can act on. Install with pip install curva-ai.