Curva

LLM classification with confidence scores you can trust.

Bring the AI key you already use. Add a few lines of code. Your AI gives one clear answer and how sure it is.

Which team should handle this email?Billing, 97% sure

Curva vs Jev on a public phishing benchmark

Curva's best run on each measure, with the model named under each chart. Jev is a hosted AI decision service.

CurvaJev

  • 18 points more accurate

    Accuracy on phishing emails

    gemini-flash-lite-latest · PhishNChips · n = 100 · 2026-10-01 · 95% range 73% to 89%

  • 10% lower error

    Calibration error: how far its confidence is from reality

    gemini-flash-lite-latest · PhishNChips · n = 100 · 2026-10-01

  • 61 ms faster

    Typical response time

    groq qwen3.8-27b · PhishNChips · n = 100 · 2026-10-01

Every number, including where Jev is ahead Jev numbers are their published figures.

Ask an AI to sort a ticket and it writes a sentence.

Your code needs one clear answer and a confidence number.

  1. You get a sentence, not a label your code can use.
  2. “Fairly sure” is not a number you can act on.
  3. Reorder the options and the answer can change.
  4. Nothing records which model answered, or how sure it was.

Which team should handle this ticket?

Without Curva

It’s probably billing, but it could be technical. I’m fairly sure, maybe 0.95.

  • no label
  • a guess
  • can’t route on it

With Curva

choicebilling confidence
{"choice": "billing", "confidence": 0.97}

Your code can act on this.

How Curva works: confidence scores you can trust.

  1. It asks twice.

    Options in one order, then reversed, and averaged. The order of options can no longer sway the answer.

    “I was charged twice, please refund me.” Which team?

    As asked
    1. billing0.78
    2. technical0.14
    3. account0.08
    Reversed
    1. account0.12
    2. technical0.18
    3. billing0.70
  2. A score for every option.

    Read from the model itself, so you see how sure it is of each one.

    1. 0.74billing
    2. 0.16technical
    3. 0.10account
  3. You send corrections.

    Tell Curva the right answer. After 30, it tunes its confidence, and keeps the change only if it helps.

    30/ 30 labels

  4. Then 90% sure means right about 90% of the time on your data.

    That is calibration: confidence that matches how often it is right. Curva is named after the curve that checks it.

    90%90% 000.50.511 confidence how often right

    0.149

    Calibration error: how far confidence is from reality, lower is better. Before tuning: 0.404.

    Gemini flash-lite, AITA, n = 100, tested on rows it did not learn from, 2026-10-01. Curve illustrative.

How Curva fits in your stack.

Your app calls Curva on your own server. Curva asks your model, reads a probability for every option, and sends back one of your labels.

How Curva fits in your stack Your app, an n8n workflow, an AI agent or any HTTP tool calls the Curva API on your own servers. Rules and the cache settle easy cases with no model call. The rest is asked twice, in both option orders, to your LLM provider. Curva reads a probability per label, averages both orders, applies the calibrator fitted from your feedback labels, and returns a typed answer. Decisions go to the audit log in an embedded SQLite database. Your app Python or TypeScript SDK n8n workflow Curva Route node AI agent decide tool Any HTTP tool Zapier, Make, curl Curva API decide, feedback, metrics Rules and cache settled with no model call Ask twice both option orders Your LLM provider OpenAI, Anthropic, Ollama Read probabilities one per label, both orders Calibrator fitted from your labels Embedded SQLite audit log, feedback labels SDK n8n node MCP HTTP API decide the rest both orders probabilities averaged typed answer audit + feedback true labels Your servers: one Curva binary
  • SDK

    Route support tickets

    Rules catch the obvious ones at no cost; unsure tickets go to a person.

  • n8n node

    Branch an n8n workflow

    The Route node gives one output per option, plus Needs review.

  • MCP

    Give an agent a decide tool

    Typed answers and an abstain flag instead of a free-text judgement.

  • SDK

    Check LLM output

    Pass, revise or block an answer before a user sees it.

  • HTTP API

    Filter phishing emails

    A probability of phishing, a risk level and a sender-mismatch flag in one call.

  • HTTP API

    Move off a hosted decision API

    Curva accepts the criteria request shape, so a client needs a new base URL and key.

Curva vs a raw LLM call, side by side.

One support ticket and one question: which team should handle it? Here is what each one hands back to your code.

ticketI was charged twice for order A-104. Please refund the duplicate!

A raw LLM call

PromptWhich team should handle this ticket? Reply with one of billing, technical, sales.

This looks like a billing issue, fairly sure.

  • A sentence. Your code has to parse it and hope it named a team that exists.
  • “Fairly sure” is not a number you can set a threshold on.
  • Reorder the options and many models change their answer.

Curva

{"decision_id": "dec_…", "answers": {
  "department": {"choice": "billing", "probabilities": {"billing": 0.97, "technical": 0.02, "sales": 0.01, "none_of_these": 0.0},
                 "confidence": 0.97, "abstain": false, "calibrated": false},
  "refund": {"noul": 0.93, "calibrated": false}}}
Probability for each team, from the response above
  • billing0.97
  • technical0.02
  • sales0.01
  • none_of_these0.0
  • choice is always one of your labels, checked before it reaches you.
  • A probability for every option, including none_of_these for “none of them fit”.
  • abstain turns true below the confidence you set (0.8 here), so a person can take it.
  • refund is a second question, answered in the same call: P(yes) is 0.93.

Nuance. The ticket and the numbers are the docs’ example; real numbers depend on the model you choose. A probability matches how often Curva is right only once that exact question is calibrated: after 30 feedback labels, and only if the calibrator beats the raw numbers on held-out labels. Until then you get the model’s own probabilities, debiased, which is what calibrated: false says. Debiasing asks in two option orders, so it costs two model calls unless you set debias: "auto".

Automate the sure answers. Send the rest to a person.

Set how sure Curva must be (min_confidence). Below that, it says it isn't sure (abstain: true) and hands the case to a person.

What you can build with Curva.

Pick one to see what Curva sends back.

Send every ticket to the team that should handle it.

Flag suspicious emails, with how sure Curva is.

A policy call on every post, text and images together.

Check an AI answer before your user sees it.

Pull values out of invoices and receipts, each with its own confidence.

Linked questions about one alert, answered in one request.

ticket

“Charged twice, refund please”

Which team should handle this ticket?

teambilling
confidence0.97
  • billing0.97
  • technical0.02
  • sales0.01
  • none_of_these0.00
  • refundP(yes)0.93

Under 0.8 Curva abstains. A person picks the team, and that answer calibrates the question.

email

Its link opens a webmail login page on port :2096, on a domain that has nothing to do with the sender.

Is this email phishing?

tactic foundsuspicious_link
  • phishingP(phishing)
  • tacticseach tactic found, with a probability
  • riskSafe · Low · Medium · High
  • sender_mismatchP(yes)

With Gemini flash-lite it got 80.8% right on a public phishing set (n = 100). Smaller models did worse, so test yours first.

post

A caption and an image, from an account 3 days old.

Does the image break the content policy?

Not sure enough? A moderator decides.

abstains below0.85
  • oknothing wrong
  • nudity
  • violence
  • spamads or scams

Up to 8 images per decision. The audit log keeps a hash of each image, never the image.

draft

request the user’s messageanswer your model’s draftcontext the source documents

Pass, revise or block this answer?

Unsure? The draft waits for a person.

abstains below0.8
  • verdictpass · revise · block
  • answers_questionNot at all · Partly · Mostly · Fully
  • groundedP(every claim is supported)
  • unsafeP(yes)
  • leaks_dataP(yes)

Ask two or three models at once. Where they disagree, confidence drops: your cue to look closer.

invoice

ACME Inc. / Invoice 2291 / 3 x Widget @ 4.10 / Shipping 2.00 / Total due 14.30 EUR

What are the vendor, invoice number, total and PO number?

total14.3
confidence0.94
  • vendorACME Inc.
  • invoice_no2291
  • total14.30.94
  • po_numbernull

The invoice has no PO number, so Curva returns null instead of inventing one.

alert

stage 0 which system is failingstage 1 how severe, and did a deploy cause itstage 2 roll back or not

Should the last deployment be rolled back now?

answered atstage 2
P(yes)0.78
  • systemdatabase · api · network
  • severityminor · major · critical
  • deployment_relatedyes or no
  • rollback0.78

Each question sees the earlier answers and how sure they were. The decision is logged once.

Multi-step LLM decision workflows, in one request.

Real decisions come in steps, and each answer depends on the one before. Curva runs the whole chain in one request, stage by stage. Each question sees the earlier answers and how sure they were.

Incident triage: an alert fires, and four questions are answered in three stages.
  1. stage 0
    • system

      Which system is failing?

      database · api · network
  2. stage 1
    • severity

      How severe is the incident?

      minor · major · critical
    • deployment_related

      Did a recent deployment cause it?

      yes or no
  3. stage 2
    • rollback

      Should the last deployment be rolled back now?

      yes or no, with thinknoul0.78stage2

The whole flow is one call

from curva import Curva, Choice, Noul, Score

d = Curva().decide(alert, {
    "system": Choice("Which system is failing?", ["database", "api", "network"]),
    "severity": Score("How severe is the incident?", ["minor", "major", "critical"]).depends("system"),
    "deployment_related": Noul("Did a recent deployment cause it?").depends("system"),
    "rollback": Noul("Should the last deployment be rolled back now?", think=True)
        .depends("severity", "deployment_related"),
})

The rollback question gets think=True: the model reasons in its own call there, while the other questions answer fast. The decision is logged once.

What else a flow can use

Rules rules
Easy cases are answered on the spot, with no model call: “mentions an invoice, so billing”.
Follow-ups when · @key
Ask a question only when an earlier answer calls for it. Skipped questions are never sent or paid for.
Cascade cascade
A cheap model answers first. Only the questions it is unsure about go to the strong one.
Council council
Two to five models at once, their probabilities blended. Where they disagree, confidence drops.
Extraction Text · Number · Integer
Read the total from an invoice, then decide “reimbursable?” with it in view, in the same request.

Nuance. A request runs at most 8 stages, and latency and cost add up across them. think is slower and costs the reasoning tokens, so keep it for the hard question. A council or a cascade means more model calls. And like the models it uses, Curva is bad at counting, arithmetic and comparing dates: compute those in code and put the result in the state.

Type-safe LLM output: answers that always fit your types.

Curva never reads a label out of a sentence. It reads the probability of each label you declared, or asks for JSON limited to them, and checks every answer before it reaches you.

Choices: only your labels, plus a way out

  • billing
  • technical
  • none_of_theseadded by default

The answer is always one of these keys. When none of your options fit, the model can say so with none_of_these instead of being forced into a wrong pick.

In TypeScript, the answer’s type is built from your options, so a typo is a compile error:

d.answers.team.choice;     // "billing" | "technical" | "none_of_these", typed from the options

Values: checked against their type and bounds

invoiceACME Inc. / Invoice 2291 / 3 x Widget @ 4.10 / Shipping 2.00 / Total due 14.30 EUR

  1. totalNumber, min=0

    14.3confidence 0.94

    A number, inside its bounds.

  2. po_numberText, nullable

    None

    Not in the document, so it says so instead of making one up.

  3. any valuewrong type or out of range

    nullconfidence 0

    Rejected. Never a silent bad value.

Text, Number and Integer questions each come back as a value with a confidence, so a cut-off works on extraction too.

Nuance. A valid type is not a right answer: a well-formed label can still be wrong, which is what the probabilities, abstain and calibration are for. none_of_these is on by default and you can turn it off; the LLM output QA recipe’s verdict has no escape option. Text, Number and Integer questions are answered in verbal mode. Keep arithmetic in code: ask for the numbers printed on the document, then add them up yourself.

Curva vs Jev: accuracy on public benchmarks.

How often each one picks the right answer, on six public test sets. We call it a win only when even the low end of Curva's range beats Jev. Otherwise it matches.

LLM classification accuracy, test set by test set

Run 2026-10-01 · Jev scores are its published ones, on different samples

  • Curva, win
  • Curva, matches
  • 95% range, the likely spread
  • Jev (published)
  1. PhishNChipsgemini-flash-lite-latest · n = 100

    80.8%Jev 62.6%win

  2. PhishNChipsqwen3.8-27b on Groq · n = 100

    72.8%Jev 62.6%win

  3. HellaSwagqwen3.8-27b on Groq · n = 100

    86.0%Jev 86.1%matches

  4. OpenBookQAgemini-flash-lite-latest · n = 100

    92.2%Jev 94.2%matches

  5. BoolQgemini-flash-lite-latest · n = 127

    88.9%Jev 89.7%matches

  6. BANKING77gemini-flash-lite-latest · n = 120

    79.8%Jev 75.3%matches

Asking an LLM directly vs asking through Curva
Asking an LLM directlyCurva
OutputA sentence you parseOne of your labels, validated
ConfidenceA self-reported guessA probability for every option. Once you send corrections, it matches how often it is right
StabilityCan change when options are reorderedAsked in both orders and averaged
When unsureIt guessesSays it is not sure (abstain: true), so a person can decide
AuditNothing to checkEvery decision logged
Runs onTheir serversYours, with any model

Jev is still ahead on some runs. Curva vs Jev, every number ↗

Jev and TypeSafe are trademarks of their respective owners. Tarkova is not affiliated with them. Jev numbers are their published figures.

Curva works with the model you already use.

We ran eight public test sets through Curva with six models from four providers. Each line is one model. The pink diamonds are Jev's published scores.

Accuracy by model, test set by test set

Run 2026-10-01 · share of right answers at each set's natural mix (Yelp: plain) · solid: n = 69 to 127 · dashed and hollow: early, n = 20

Show

Jev, published scoreOn its own samples, 77 to 2,000 items

  • PhishNChips: both free models beat Jev's 62.6%. Gemini scores 80.8% (73% to 89%), qwen3.8-27b on Groq 72.8% (64% to 82%), n = 100 each.
  • BANKING77, OpenBookQA, CommonsenseQA, HellaSwag and BoolQ: the best free model matches Jev, within the margin.
  • AITA: Jev is ahead, at 75.4% against Gemini's 67.2% (within the margin) and Groq's 55.7%.
Every number in a table
Accuracy at each set's natural label mix (Yelp: plain accuracy), with the 95% range and n. Run 2026-10-01. Paid models are early, n = 20 per set.
ModelPhishNChipsBANKING77OpenBookQACommonsenseQAHellaSwagAITABoolQYelp
gemini-flash-lite-latest80.8%73% to 89%n = 10079.8%73% to 87%n = 12092.2%87% to 97%n = 10082.0%74% to 90%n = 10078.6%70% to 87%n = 9867.2%58% to 76%n = 10088.9%83% to 94%n = 12754.0%44% to 64%n = 100
qwen3.8-27b on Groq72.8%64% to 82%n = 10074.7%64% to 85%n = 6985.0%78% to 92%n = 10084.0%77% to 91%n = 10086.0%79% to 93%n = 10055.7%46% to 65%n = 100not runnot run
Claude Haiku 4.5 (early)75.0%56% to 94%n = 2065.0%44% to 86%n = 2094.5%84% to 100%n = 2075.0%56% to 94%n = 2085.0%69% to 100%n = 2061.2%40% to 83%n = 2090.0%77% to 100%n = 2050.0%28% to 72%n = 20
gpt-4.1-mini (early)70.0%50% to 90%n = 2075.0%56% to 94%n = 2091.7%80% to 100%n = 2080.0%62% to 98%n = 2080.0%62% to 98%n = 2041.0%19% to 63%n = 2090.0%77% to 100%n = 2050.0%28% to 72%n = 20
gpt-4o-mini (early)50.0%28% to 72%n = 2060.0%39% to 81%n = 2086.4%71% to 100%n = 2090.0%77% to 100%n = 2055.0%33% to 77%n = 2040.9%19% to 62%n = 2095.0%85% to 100%n = 2055.0%33% to 77%n = 20
gpt-4.1-nano (early)55.0%33% to 77%n = 2050.0%28% to 72%n = 2077.0%59% to 95%n = 2075.0%56% to 94%n = 2060.0%39% to 81%n = 2040.9%19% to 62%n = 2075.0%56% to 94%n = 2050.0%28% to 72%n = 20
Jev (published)62.6%75.3%94.2%88.1%86.1%75.4%89.7%none

Accuracy and speed on PhishNChips

Median time per decision (log scale) · run 2026-10-01 · vertical lines are the 95% range

Faster and more accurate than Jev 20% 40% 60% 80% 100% 200 ms 500 ms 1,000 ms gemini-flash-lite-latest 80.8% 1,146 msgemini-flash-lite-latest · PhishNChips
80.8% (95% range 73% to 89%), n = 100
Median 1,146 ms per decision
qwen3.8-27b on Groq 72.8% 178 msqwen3.8-27b on Groq · PhishNChips
72.8% (95% range 64% to 82%), n = 100
Median 178 ms per decision
Claude Haiku 4.5 75.0% 958 msClaude Haiku 4.5 · PhishNChips
75.0% (95% range 56% to 94%), n = 20, early
Median 958 ms per decision
gpt-4.1-mini 70.0% 720 msgpt-4.1-mini · PhishNChips
70.0% (95% range 50% to 90%), n = 20, early
Median 720 ms per decision
gpt-4o-mini 50.0% 805 msgpt-4o-mini · PhishNChips
50.0% (95% range 28% to 72%), n = 20, early
Median 805 ms per decision
gpt-4.1-nano 55.0% 750 msgpt-4.1-nano · PhishNChips
55.0% (95% range 33% to 77%), n = 20, early
Median 750 ms per decision
Jev 62.6% 239 msJev · PhishNChips
62.6% on 2,000 emails (published)
Median 239 ms, measured by a third party

qwen3.8-27b on Groq is faster and more accurate than Jev on this set: 72.8% in a median 178 ms, against 62.6% in 239 ms (n = 100). Gemini is the most accurate and the slowest. Every other model is slower than Jev.

Why you can trust these numbers

  1. Public test sets. All eight sets are public, so anyone can check what was asked: PhishNChips, BANKING77, OpenBookQA, CommonsenseQA, HellaSwag, AITA, BoolQ and Yelp reviews.
  2. A fixed sample. Rows are drawn with a fixed random seed, not picked by hand, and balanced across the answers.
  3. The n and the range, every time. Each score carries its sample size and its 95% range. Accuracy is reweighted to each set's real mix of answers, the way a published score is measured. Paid models ran 20 rows per set, so we mark them early: at n = 20 the range is about 20 points either way.
  4. Jev's published numbers. Jev's scores are the ones published for it, on different samples (77 to 2,000 items). Compare with care. We call it a win only when the low end of Curva's range is above Jev's score.
  5. Where Jev is ahead, too. Jev leads on AITA, on OpenBookQA against Groq, on raw calibration for BoolQ and the multiple-choice sets, and on speed against every model but Groq. Curva vs Jev, every number ↗

Cost per 1,000 decisions is what each provider charged across 160 decisions per paid model, 20 on each of the eight sets (2026-09-30). Gemini and Groq ran on their free tiers, within the daily limits. Two more free models ran on PhishNChips only (gpt-oss-20b, n = 36; Nemotron 3 Super, n = 20) and are not charted. Model and company names and logos belong to their owners.

Add Curva to your app in three lines.

Use the AI API key you already have: OpenAI, Anthropic, Gemini, Groq, OpenRouter, or a local model. Add a few lines of code, and your AI makes quick decisions you can trust. Curva itself is free.

pip install curva-ai
export OPENROUTER_API_KEY=sk-or-v1-...     # or any one provider key
import curva
d = curva.decide("I was charged twice, please refund me",
                 {"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total)           # e.g. billing True None
teambilling refundTrue totalNone

It also works from TypeScript, n8n, MCP and Docker. See the docs ↗

Your AI key. A few lines of code. Clear answers you can trust.