How Curva works: from state to calibrated answer

The path of one Curva request: rules, cache, fenced prompt, two option orders, logprobs or verbal reading, calibration, coverage set, abstain and audit.

How Curva works, in one sentence: you send data (the state) and typed questions, Curva settles what it can without a model, asks the model in two option orders and reads a probability for every label, adjusts those probabilities with a calibrator fitted on your feedback when that helps, and returns a typed answer that is written to an audit log. Every step leaves a mark in the response: mode, debiased, cached, calibrated, set, abstain, stage, rule. This post walks the path of one request, step by step, and names the field each step adds.

Curva is a small server and SDKs. It is one binary with an embedded SQLite database, free to use under the Curva Free License, and it calls whichever LLM you configure.

How Curva works: the request path at a glance

flowchart TD
  A["state + typed questions"] --> B{"rule matches or when fails?"}
  B -- yes --> Z["answer or skipped, no model call"]
  B -- no --> C{"identical decision cached?"}
  C -- yes --> Y["cached answer, about 1 ms"]
  C -- no --> D["prompt: state fenced as data"]
  D --> E["model asked in original and reversed option order"]
  E --> F["probabilities read: logprobs or verbal"]
  F --> G["average both orders"]
  G --> H["calibrate, only if it beats raw"]
  H --> I["coverage set and abstain flag"]
  I --> J["typed answer + audit log"]
  J --> K["your feedback refits the calibrator"]
Figure 1. How Curva works: the path of one request Questions answered by rules, skipped by when, or found in the cache never reach the model.

State and typed questions in, up to 64 per request

The state is the data to judge. It is any JSON: a string, an object or a list, up to 150,000 characters, plus up to 8 images for vision models. The questions are typed, up to 64 per request, and all are answered in one request.

TypeAsksAnswer
ChoicePick one of 2 to 255 optionschoice, a probability per option, confidence
ScorePlace on a scale of 2 to 20 levelsscore, the expected level (0 = lowest), a probability per level
NoulYes or nonoul = P(yes)
MultiPick any subset of 1 to 20 optionsselected, an independent probability per option
Text, Number, IntegerRead a valuevalue, confidence

Answers are only ever mapped onto the labels you declared. Nothing is parsed from free text, so there is nothing to validate on your side.

Settled before the model: rules, when and the cache

Three checks can answer a question without a model call.

**Rules.** A question with rules ("if the ticket contains invoice, answer billing") is answered on the spot when a rule matches. Up to 32 per question, first match wins. The answer has confidence: 1.0, calibrated: false and rule, the index of the rule that fired.

**When.** A question whose when doesn't match the state is skipped. It comes back as {"skipped": true} and is never sent, stored or paid for. If every question is answered by a rule or skipped, no model is called and the decision costs $0.

**The decision cache.** Calls run at temperature 0, so an identical decision (same model, state, questions, config, project, privacy and think) is answered from memory in about a millisecond. The response says cached: true, with zero latency and cost.

The prompt: state fenced as data, both option orders

Curva builds a prompt with the state wrapped in a fenced block, and tells the model it is data, not instructions. Any closing fence inside the data is escaped, and the fence is matched case-insensitively. That is Curva's defence against prompt injection in the state, and an adversarial eval set tracks it.

Then it asks. By default, every question is asked twice, concurrently: once with the options in your order and once reversed. Models tend to favour options by position, and averaging the two orders cancels that. Few-shot examples are remapped so they stay correct in the reversed call. Options named A, B, C keep their order and take one call. The response's debiased field says whether both orders were asked.

The model can be one model, a fallback chain, or a council, cascade or race of several. A Choice also gets an extra option, none_of_these, so the model is never forced into a wrong pick.

Reading probabilities: logprobs, verbal, auto

Curva reads probabilities. It never parses a sentence.

  • **logprobs**: the model answers with one label token per question, and Curva reads each label's probability from the token log-probabilities. If your labels hold less than half the probability at that position, that question is retried on its own.
  • **verbal**: the model returns a JSON object limited to your labels, with a probability for each.
  • **auto** (the default): logprobs when the model returns them, verbal otherwise.

The response's mode says which one answered. Choices with more than 20 options, and requests with a Text, Number or Integer question, always use verbal mode, because providers return at most 20 logprobs and extraction has no label token to read. A Noul is asked as a normalised two-option choice, so P(x) and P(not x) stay consistent.

After both orders answer, Curva averages the probabilities option by option.

After the model: calibrate, coverage set, abstain

Three steps turn raw probabilities into something you can automate on.

**Calibrate.** Once a question has 30 feedback labels in a project, Curva fits a small calibrator: temperature scaling for Choice and Score, bias scaling (temperature plus a per-answer offset) for Choice and Score with up to 20 options when it beats temperature alone, and Platt scaling for Noul. It keeps the calibrator only if it beats the raw probabilities on held-out labels, in both log-loss and Brier score, by more than its own noise. Then answers carry calibrated: true. Otherwise they stay raw.

**Coverage set.** With coverage on a Choice, Multi or Noul, and 30 labels, the answer gains guaranteed: true and a set of options. Curva uses split conformal prediction on your feedback labels, so the set contains the true answer at least coverage of the time, as long as new inputs look like the labeled ones.

**Abstain.** With min_confidence on a question, answers below it carry abstain: true, your signal to hand the item to a person.

Here is a real response from the docs, before any of those three were set:

json
{
  "id": "dec_19294a3c1f2000000",
  "model": "inclusionai/ling-3.0-flash-fin:free",
  "mode": "logprobs",
  "latency_ms": 1144,
  "cost_usd": 0.0,
  "cached": false,
  "answers": {
    "department":       { "choice": "billing", "probabilities": { "billing": 0.9999, "technical": 0.0, "sales": 0.0001 }, "confidence": 0.9999 },
    "frustration":      { "score": 0.65, "probabilities": [0.36, 0.62, 0.02], "confidence": 0.62 },
    "refund_requested": { "noul": 0.999 }
  }
}

The Score shows why the expected level matters: probabilities of 0.36, 0.62 and 0.02 give a score of 0.65, between the first and second level, instead of rounding the doubt away.

Stages for depends_on

Some questions only make sense after others. Give a question depends_on and Curva runs the request as a graph, stage by stage, with one model call per stage and at most 8 stages. Each dependent question sees the earlier answers and their probabilities in a fenced block. Every answer then carries stage, its 0-based stage number. when and rules can read an earlier answer with @key, so a follow-up is asked only when an earlier answer calls for it. A skipped dependency skips its dependents.

A question can also set think: true to reason in a verbal call of its own, next to the fast call for the rest of its stage.

Audit log and the feedback loop

Every decision is written to the audit log: its id, project, API key id, a salted hash of the state (never the state itself), model, mode, config, privacy setting, answers, latency, cost and whether it was cached. Images are logged as a salted SHA-256, never stored. GET /v1/audit pages through it, newest first.

The loop closes when you learn the truth. POST /v1/feedback records the right answer for one question of a past decision, and Curva refits that question's calibrator on every new label. Calibrators belong to a fingerprint of the question's type, wording and options, and to a project. Rewording starts fresh.

One binary, one SQLite file

All of this runs in one process. The curva binary is the CLI and the HTTP server, and its database (decisions, feedback, calibrators, audit log, keys) is one SQLite file. A Docker image of about 62 MB is published for amd64 and arm64.

Curva's own overhead is small next to the model's. With a fake model on a laptop, the docs report p50 0.59 ms at 500 requests per second, and about 8,000 requests per second saturated. With real remote models, a decision takes about 0.5 to 3 s on free providers. In practice the provider is the limit.

Next steps

The docs explain each part in questions and answers, debiasing, escape and abstain and probabilities and calibration. For the product view, read what is Curva. For the mechanisms in depth, see LLM position bias, logprobs vs verbal confidence and conformal prediction for LLMs. Install with pip install curva-ai.