Abstain or a coverage set? Selective classification for LLMs

Abstain checks the top answer; a conformal set checks every option. A decision table by question type and the routing rule that combines both.

Selective classification vs conformal prediction is a choice between two ways of handling an unsure LLM answer. Selective classification lets the classifier abstain: if the top answer's confidence is below a threshold, it hands the item to a person. Conformal prediction keeps every answer but attaches a set of options that contains the true one at a promised rate. In Curva the first is min_confidence, which sets abstain: true, and the second is coverage, which adds a set. Abstain looks at the top answer; a coverage set looks at every option. They answer different questions, they work on different question types, and the strongest routing rule uses both.

Two questions: is the top answer good enough, which options must I keep

Every routing decision on an LLM answer comes down to one of two questions.

  • **Is the top answer good enough to act on?** That is selective classification. You set a bar, and answers below it abstain. In Curva: min_confidence on the question, and abstain: true on the answer.
  • **Which options must I keep to be right often enough?** That is conformal prediction. You set a target rate, and each answer comes with the smallest set of options that meets it. In Curva: coverage on the question, and set plus guaranteed on the answer.

A word on vocabulary, because the two fields collide. In the selective classification literature, "coverage" usually means the share of inputs the classifier answers instead of abstaining. In Curva, coverage is the conformal target: the probability that the set contains the true answer. Curva's own name for the share answered is automated, in the calibration report.

Selective classification vs conformal prediction: support by question type

The two are not available on the same question types:

Question type`min_confidence` (abstain)`coverage` (set)
Choiceyesyes
Scoreyesno
Noulnoyes, a subset of ["true", "false"]
Multinoaccepted, but guaranteed: false for now
Text, Number, Integeryesno

Two gaps are worth knowing. A Noul has no min_confidence: its answer is one number, P(yes), so you route it with your own thresholds, treating the middle band as unsure, or you set coverage and read the set. A set of ["true", "false"] means the model can't tell. And a Multi accepts coverage, but feedback labels a Multi as a whole rather than option by option, so its sets carry no guarantee yet. Setting coverage on a type that doesn't support it gets 422.

What each needs: a calibrated threshold vs 30 representative labels

Both tools work from day one, but neither is fully meaningful until it has labels.

**Abstain needs calibration.** min_confidence compares the answer's confidence with your bar. Before the question is calibrated, that confidence is the model's own number, debiased across two option orders but not checked against your data. A bar of 0.8 is then a guess about the model. Once your feedback has fitted a calibrator (after 30 labels, and only when it beats the raw probabilities on held-out labels), the threshold means what it says.

**A coverage set needs 30 representative labels.** The set is built by split conformal prediction from the question's own feedback labels. Before 30 labels, the answer has guaranteed: false and no set. After 30, the guarantee holds whatever the model, as long as the labels are a fair sample of your traffic. Calibration helps by making sets smaller, but the guarantee doesn't depend on it.

AbstainCoverage set
Looks atthe top answerevery option
Promisenone until calibrated; then the threshold means what it saysthe true answer is in the set at least coverage of the time
Needsa calibrator for the question30 labels that look like your traffic
Fails quietly whenthe model is overconfident and uncalibratedlabels are not representative or traffic shifts

Combined rule: a one-option set and abstain false

The two checks catch different failures, so combine them. Curva's docs give a common routing rule: automate when the set has one option and abstain is false, and send everything else to review.

python
q = Choice("Which team?", ["billing", "technical", "sales"], min_confidence=0.8, coverage=0.95)
d = curva.decide(ticket, {"team": q}, project="support")
a = d["team"]
if a.guaranteed and len(a.set) == 1 and not a.abstain:
    route(ticket, a.choice)
else:
    review(ticket, candidates=a.set or [a.choice])

Why both? A one-option set says the model is sure enough to meet the coverage guarantee alone. abstain: false says the top answer also clears your own bar. Either check alone can let an answer through that the other would stop. And the guaranteed check guards the start: until the question has 30 labels, there is no set and everything goes to review.

Showing the set to a reviewer as candidates

The else branch above doesn't just send the item to a person. It sends the candidates. That is where a coverage set earns its keep even when it can't automate.

SetMeaningWhat the reviewer sees
One optionsure enough to meet the guarantee aloneusually automated; shown if abstain was true
Several optionsthe truth is one of these, at the promised ratea short list to choose from
Every optionthe model can't narrow it downthe full list, so a fresh read

A reviewer choosing between two likely teams works faster than one starting from the full list. When the reviewer decides, send the answer back with client.feedback(decision_id, "team", "technical"). Each review adds a label that feeds both calibration and the conformal threshold.

Metrics to watch for each

Each tool has its own health signal.

**For abstain**, watch how much you automate and how well it goes. The calibration report gives automated, the share of answers at 0.9 or above, and accuracy_when_automated, how often those were right, before and after calibration. The server's curva_abstains_total counter, and the abstain rate on the dashboard's overview, show how much traffic reaches people.

**For a coverage set**, watch two things. First, the size of the sets: the share with one option is how much you can automate under the guarantee. Second, whether the sets still cover. Label a random sample of recent decisions and count how often the true answer was inside the set. At 0.95 coverage, expect about one miss in twenty; many more suggests the labels no longer match your traffic, and the weekly drift report is the place to check.

Neither replaces the other. Abstain is the simpler tool and works on Scores and extraction. Coverage gives a promise you can state in a contract with your own team. Use abstain everywhere it applies, and add coverage where a stated rate matters.

Next steps

The docs describe both in guaranteed accuracy and debiasing, escape and abstain. Go deeper on each side with LLM confidence threshold and abstain and conformal prediction for LLM classification. For the accuracy you get at each threshold, read selective accuracy for LLMs, and for the full picture, LLM classification confidence scores you can act on.