Python LLM classification with confidence and feedback

A start-to-finish Python tutorial: typed LLM labels with a probability each, abstain, feedback, calibration reports, async batches and errors.

To do Python LLM classification with a confidence you can act on, install curva-ai, declare your labels as typed questions, and call decide. You get one of your labels back with a probability for every option, never free text to parse. Send the true answer when you learn it, and after 30 labels Curva calibrates those probabilities on your data. This tutorial covers the whole loop: install, decide, abstain, feedback, the calibration report, async batches and error handling.

Most Python classification code built on an LLM looks the same today: a prompt that says "reply with one of: billing, technical, sales", a call, and a parser that hopes the reply is one of those words. Then someone adds "and how confident are you?", and the model writes a high number on almost everything, right or wrong. The steps below replace that pattern.

Install curva-ai

bash
pip install curva-ai
export OPENROUTER_API_KEY=sk-or-v1-...     # any OpenRouter key; free models work

The package includes the Curva server binary for Linux, macOS and Windows, so there is nothing else to install. It needs Python 3.9 or newer and uses only the standard library. Other providers work too: without OPENROUTER_API_KEY, the first of OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, GROQ_API_KEY and a few more picks a small, fast model of that provider. Set CURVA_MODEL to choose one yourself. Curva is free to use; you pay only your model provider, and free models work.

Python LLM classification in three lines

python
import curva
d = curva.decide("I was charged twice, please refund me",
                 {"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total)           # billing True None

The first curva.decide starts a private server for this process and reuses it. The shorthand maps Python values to question types:

You writeYou get
["billing", "technical"] or {"billing": "payments"}Choice
"Asks for a refund?" (ends in ?)Noul (yes or no)
bool, float, int, strNoul, Number, Integer, Text, named after the key

d.team is the plain answer. d["team"] is the full one, with confidence and probabilities. d.to_dict() gives all plain answers as a dict. For a yes/no question, the plain answer is True when P(yes) is at least 0.5; read d["refund"].noul when you want the probability itself. The shorthand is good for scripts and notebooks. For anything you will run in production, switch to explicit question objects, because they let you set thresholds, examples and coverage per question.

The full client with typed questions

For real code, use explicit question objects and a client:

python
import curva
from curva import Choice, Score, Noul

client = curva.local()      # or curva.Curva("http://your-server:7777")

QUESTIONS = {
    "team": Choice("Which team should handle this?",
                   {"billing": "payments, refunds", "technical": "bugs", "sales": "pricing"},
                   min_confidence=0.8),
    "frustration": Score("How frustrated is the customer?", ["calm", "annoyed", "angry"]),
    "refund": Noul("The customer explicitly asks for a refund"),
}

d = client.decide({"ticket": "I was charged twice for order A-104. Please refund the duplicate!"},
                  QUESTIONS, project="support")

print(d["team"].choice, d["team"].confidence)   # e.g. billing 0.9999
print(d["frustration"].score)                   # e.g. 0.65
print(d["refund"].noul)                         # e.g. 0.999
print(d.mode, d.latency_ms, d.cost_usd, d.cached)

A few details worth knowing:

  • state is any JSON value. It is fenced as data in the prompt, never treated as instructions.
  • All questions are answered in one request.
  • The Choice gets an extra option, none_of_these, so the model is never forced into a wrong pick.
  • score is the expected level: 0.65 sits between "calm" (level 0) and "annoyed" (level 1).
  • project is a calibration namespace. Use one per use case.
  • curva.local() keeps its database in ~/.curva/curva.db, so what it learns survives restarts.
  • A Score answer knows its level names: d["frustration"].level is the most likely one.

Abstain: send unsure answers to a person

Because the Choice set min_confidence=0.8, answers below 0.8 come back with abstain=True:

python
answer = d["team"]
if answer.abstain:
    team = ask_a_human(ticket)
    client.feedback(d.id, "team", team)      # the person's answer teaches Curva
else:
    route(ticket, answer.choice)

abstain is only set on questions that have min_confidence. Until the question is calibrated, 0.8 is a guess about the model. Once calibrated, it means what it says. The thinking behind the threshold is in LLM classification confidence scores you can act on.

Feedback and the calibration report

Keep d.id next to whatever you did with the answer. When you learn the truth (an agent closes the ticket, a reviewer fixes a label), send it:

python
client.feedback(decision_id, "team", "billing")     # Choice: the option key
client.feedback(decision_id, "frustration", 2)      # Score: the level index
client.feedback(decision_id, "refund", True)        # Noul: True or False

Sending feedback again for the same decision and question replaces the earlier label. After 30 labels for a question in a project, Curva fits a calibrator. It keeps it only when it clearly beats the raw probabilities on held-out labels in both log-loss and Brier score; then answers come back with calibrated=True. Otherwise the raw probabilities stay, because the model is already well calibrated. Check what happened:

python
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])

after is held out: each half of the labels is calibrated by a fit on the other half, so the number isn't flattered by testing on the data it was fitted to. Also send feedback for a sample of confident answers, not only the abstained ones, so calibration covers the whole range.

One rule to remember: calibration belongs to the exact wording of a question. Settle the wording before you collect labels.

Many records at once with AsyncCurva

For a batch, run the async client against a server:

python
import asyncio
from curva import AsyncCurva

async def classify(rows):
    async with AsyncCurva() as client:          # CURVA_BASE_URL, CURVA_API_KEY
        results = await client.decide_many(
            [({"ticket": r["text"]}, QUESTIONS) for r in rows],
            concurrency=16, project="support", return_exceptions=True,
        )
    return [
        {"id": r["id"], "error": str(d)} if isinstance(d, Exception)
        else {"id": r["id"], "decision_id": d.id, "team": d["team"].choice,
              "needs_human": bool(d["team"].abstain)}
        for r, d in zip(rows, results)
    ]

Results keep the input order. return_exceptions=True puts each failure in its slot, so one bad row doesn't fail the batch. concurrency defaults to 8. For files on disk, curva map on the command line is the other route: it resumes after a stop.

Error handling with CurvaError

The client retries 429 and 5xx responses (three times by default) and honours Retry-After. Anything it doesn't retry raises a CurvaError with status, type and message. Catch a subclass when you care about one case:

python
from curva import CurvaError, RateLimitError, InvalidRequestError

try:
    d = client.decide(state, QUESTIONS)
except RateLimitError as e:
    wait(e.retry_after)
except InvalidRequestError as e:
    log.error("bad question: %s", e.message)    # a 422 names the question
except CurvaError as e:
    log.error("%s %s", e.status, e.type)

Status 0 means the server wasn't reachable.

Move to a shared server

curva.local() is right for scripts and notebooks. For a service, run one server and point clients at it:

bash
curva serve --addr 127.0.0.1:7777 --db curva.db
curva keys create --name my-service
python
from curva import Curva
client = Curva()        # reads CURVA_BASE_URL and CURVA_API_KEY

A shared server keeps one audit log, one set of calibrators per project, and one decision cache. Identical requests come back from the cache in about a millisecond, at no cost.

What not to ask the model

Don't ask a model to count, add up or compare dates. Do that in Python and put the result in the state. Ask Curva the judgement questions: which team, how urgent, is this a refund request.

Next steps

Read the Python SDK reference and the feedback guide in the docs. To see why Curva asks every question in two orders, read LLM position bias. For a full worked use case, try phishing detection with an LLM.