Curva concepts
The ideas behind Curva: confidence scores, calibration, position bias and letting the model abstain. 14 articles.
Subscribe with RSSStop parsing LLM text for decisions
Decisions need types and probabilities, not prose. Why parsing a model's sentence fails quietly, and what a typed, calibrated decision gives you instead.
More articles
-
The fair coin test: can an LLM say 50%?
A coin flip has one honest probability. One published probe got 0.92 from a typed-decision model; Curva gave 0.494. What the test shows and its limits.
-
Expected calibration error (ECE) for LLM classifiers
How expected calibration error is computed, worked by hand on 100 answers, and why an LLM can match on accuracy yet lose on ECE, with measured numbers.
-
Conformal prediction sets for LLM classification
Set coverage=0.95 and each answer carries a set of labels holding the right one at least 95% of the time. How split conformal works and what it needs.
-
When conformal guarantees fail: exchangeability for LLMs
Conformal prediction sets need labels that look like your traffic. How review-only labels and shifting inputs break coverage, and how to catch it early.
-
Bias scaling: fix an LLM that always picks one answer
When an LLM favours one class, temperature scaling cannot reorder answers. Bias scaling adds a per-answer offset; measured on AITA and Yelp, held out.
-
Abstain or a coverage set? Selective classification for LLMs
Abstain checks the top answer; a conformal set checks every option. A decision table by question type and the routing rule that combines both.
-
LLM position bias and how order debiasing cancels it
LLMs favour options by their position in a list. How asking in two orders and averaging cancels it, what it costs, and when debias auto skips a call.
-
LLM classification confidence scores you can act on
How to get confidence scores from LLM classification that mean something: a probability per label, debiasing, calibration on your labels and abstain.