Curva vs Jev: benchmarks, calibration, speed and cost

Both make an AI pick from your options and tell you how sure it is. Here is how they compare on public test sets, with the model, sample size and date next to every number.

  • Curva is a free server you run yourself. It gets typed answers with a probability for every option from any LLM you choose, and tunes that confidence from your corrections.

  • Jev is TypeSafe AI's "System One" model: one trained, closed model behind a hosted API that returns typed answers with probabilities.

With Curva, you use the AI API key you already have (OpenAI, Anthropic, Gemini, Groq, OpenRouter, or a local model), add a few lines of code, and your AI makes quick decisions you can trust. Curva itself is free.

In short: Curva is ahead on PhishNChips accuracy, calibration and speed (n = 100). It matches Jev on five more test sets. Jev is ahead on response time with most models, on calibration error for multiple-choice sets, on AITA, and on cost per decision with paid models.

Try CurvaWhere Jev is ahead

Benchmarks run 2026-10-01. Page updated .

Where Curva beats Jev

Measured wins only: Curva's whole 95% range sits above Jev's published number, or the benchmark record marks it a win. All on PhishNChips, a public phishing email set, n = 100, run 2026-10-01.

  • 80.8%

    Jev 62.6%

    Accuracy: how often it picks the right answer

    Gemini flash-lite · PhishNChips · n = 100 · 2026-10-01 · 95% range 73% to 89%

  • 0.138

    Jev 0.154

    Calibration error: how far confidence is from reality, lower is better

    Gemini flash-lite · PhishNChips · n = 100 · 2026-10-01 · unchanged after tuning

  • 72.8%

    Jev 62.6%

    Accuracy, on a fast free model

    Groq qwen3.8-27b · PhishNChips · n = 100 · 2026-10-01 · 95% range 64% to 82%

  • 178 ms

    Jev 239 ms

    Typical response time (median)

    Groq qwen3.8-27b · PhishNChips · n = 100 · 2026-10-01

Also ahead, too small or narrow to lead with: on AITA, Groq's typical response time is 265 ms vs Jev's 390 ms (n = 100), but its AITA accuracy is behind.

Accuracy on every test set

The dot is Curva's accuracy; the bar is its 95% range, where the true score likely sits at this sample size. The outlined mark is Jev's published number.

  • Curva, win
  • Curva, other results
  • 95% range
  • Jev (published)
  1. PhishNChipsGemini flash-lite · n = 100

    80.8%Jev 62.6%Curva wins

  2. PhishNChipsGroq qwen3.8-27b · n = 100

    72.8%Jev 62.6%Curva wins

  3. BANKING77Gemini flash-lite · n = 120

    79.8%Jev 75.3%matches

  4. BANKING77Groq qwen3.8-27b · n = 69

    74.7%Jev 75.3%within margin

  5. BoolQGemini flash-lite · n = 127

    88.9%Jev 89.7%matches

  6. OpenBookQAGemini flash-lite · n = 100

    92.2%Jev 94.2%matches

  7. OpenBookQAGroq qwen3.8-27b · n = 100

    85.0%Jev 94.2%Jev ahead

  8. CommonsenseQAGroq qwen3.8-27b · n = 100

    84.0%Jev 88.1%matches

  9. CommonsenseQAGemini flash-lite · n = 100

    82.0%Jev 88.1%within margin

  10. HellaSwagGroq qwen3.8-27b · n = 100

    86.0%Jev 86.1%matches

  11. HellaSwagGemini flash-lite · n = 98

    78.6%Jev 86.1%within margin

  12. AITAGemini flash-lite · n = 100

    67.2%Jev 75.4%within margin

  13. AITAGroq qwen3.8-27b · n = 100

    55.7%Jev 75.4%Jev ahead

Curva runs 2026-10-01, free models, n = 69 to 127 per row. "Matches": Jev's number is inside Curva's range. "Within margin": also inside, but Curva's score is clearly lower, so too close to call at this sample size.

Every number, side by side

Calibration error measures how far stated confidence is from how often the answer is right; lower is better. "After tuning" is measured on rows the tuning did not see, which is what you get once your corrections flow in.

Curva vs Jev, all free-model results, run 2026-10-01
Test setCurva modelnCurva accuracy (95% range)Jev accuracyResultCurva calibration error, before / after tuningJev calibration errorCurva typical time (ms)Jev typical time (ms)
PhishNChipsGemini flash-lite10080.8% (73% to 89%)62.6%Curva wins0.138 / 0.1380.1541,146239
PhishNChipsGroq qwen3.8-27b10072.8% (64% to 82%)62.6%Curva wins0.283 / 0.2090.154178239
BANKING77Gemini flash-lite12079.8% (73% to 87%)75.3%matches0.097 / 0.097not published1,362not published
BANKING77Groq qwen3.8-27b6974.7% (64% to 85%)75.3%within margin0.252 / 0.157not published351not published
BoolQGemini flash-lite12788.9% (83% to 94%)89.7%matches0.081 / 0.0810.0381,071not published
OpenBookQAGemini flash-lite10092.2% (87% to 97%)94.2%matches0.056 / 0.0560.0241,013not published
OpenBookQAGroq qwen3.8-27b10085.0% (78% to 92%)94.2%Jev ahead0.051 / 0.0510.024207not published
CommonsenseQAGroq qwen3.8-27b10084.0% (77% to 91%)88.1%matches0.066 / 0.0660.032230not published
CommonsenseQAGemini flash-lite10082.0% (74% to 90%)88.1%within margin0.118 / 0.1180.0321,263not published
HellaSwagGroq qwen3.8-27b10086.0% (79% to 93%)86.1%matches0.085 / 0.0850.029209not published
HellaSwagGemini flash-lite9878.6% (70% to 87%)86.1%within margin0.032 / 0.0320.0291,173not published
AITAGemini flash-lite10067.2% (58% to 76%)75.4%within margin0.404 / 0.149not published1,318390
AITAGroq qwen3.8-27b10055.7% (46% to 65%)75.4%Jev ahead0.475 / 0.221not published265390

On PhishNChips, Gemini's calibration error (0.138) and Groq's response time (178 ms) are wins; Groq's calibration error (0.209 after tuning) is behind Jev's 0.154. Jev publishes no calibration error for AITA; its AITA Brier score is 0.369 against Curva's 0.878 (Gemini) and 1.042 (Groq). Paid models (gpt-4.1-nano, gpt-4o-mini, gpt-4.1-mini, Claude Haiku 4.5) ran only 20 rows per set, too early to compare, so they are left out of this table.

Features: Curva vs Jev

What each one offers, from Curva's docs and Jev's public sources
AreaCurvaJev
The modelAny OpenAI-compatible model: OpenAI, Anthropic, Gemini, Groq, DeepSeek, Mistral, OpenRouter, or local Ollama, vLLM, LM Studio, llama.cpp. Mix them in one request.One closed model, Jev. Architecture, weights and paper undisclosed. TypeSafe launch post ↗
Where it runsYour servers: one binary with SQLite, or Docker. Your data stays with you.TypeSafe's hosted API, served from the US West Coast. TypeSafe launch post ↗
PriceFree to use, also commercially. You pay only your model provider, and free models work.$0.042 per 1M input tokens, output free (self-reported). TypeSafe launch post ↗
Answer typesChoice (2 to 255 options), Score (2 to 20 levels), yes/no, multi-select, text, number, integer.Choice (up to 255 options), Score (2 to 10 levels), yes/no probability. Jev docs ↗
ImagesUp to 8 per decision, with vision models.Text only. Jev docs ↗
When unsureA "none of these" option by default. Below your threshold it says it is not sure, so a person decides. Or a short list that holds the right answer 95% of the time.No "none of these" option. In one test, 7 of 53 answers reached 0.9 confidence. AgentConn review ↗
Confidence on your dataTuned from your corrections, per question. Kept only when it helps on rows it did not learn from.Fixed at training (self-reported). In one test a fair coin came out 0.92 heads. Alex Molas: "Jev can't be calibrated" ↗
Linked questionsOne question can depend on another (depends_on, when), plus rules and per-question reasoning.Not in Jev's docs. Jev docs ↗
Several modelsFallback, council, cascade and race across models.One model. TypeSafe launch post ↗
AuditAn audit log for every decision, and which input fields moved the answer.No reason returned with a decision. AgentConn review ↗
VersionsPinned configs. You move curva-latest yourself.jev-latest moves to new releases; a pinned version is available. Jev docs ↗
Typical speed178 ms on Groq (PhishNChips, n = 100). 720 ms to 13,373 ms on other models. Repeat decisions come from a cache at $0.70 to 500 ms (self-reported). 239 ms measured on PhishNChips. jev-phishing-bench (PhishNChips, 2,000 emails) ↗
Input sizeUp to 150,000 characters of input. The total limit is your model's.64k tokens; input plus the longest question within 32k. Jev docs ↗
SDKsPython (sync and async), TypeScript, n8n node, MCP server, OpenAPI.Python and TypeScript. Jev docs ↗

Cost per 1,000 decisions

Curva is free to use and runs on your servers, so there is no per-call fee to Tarkova. You pay your model provider. Each decision is two model calls by default (the options are asked in both orders).

Curva, measured 2026-09-30: 160 decisions per model, 20 on each of 8 public sets
ModelPer 1,000 decisionsNote
Gemini flash-lite$0free tier, about 250 decisions a day
Groq qwen3.8-27b$0free tier, about 180 decisions a day
gpt-4.1-nano$0.078$0.055 on PhishNChips
gpt-4o-mini$0.118$0.083 on PhishNChips
gpt-4.1-mini$0.314$0.223 on PhishNChips
Claude Haiku 4.5$1.64$1.29 on PhishNChips
Any model, repeat decision$0served from the cache
Jev, published
FigureValueSource
List price$0.042 per 1M input tokens, output free (self-reported)TypeSafe launch post ↗
Per 1,000 decisions$0.025 (third-party, 4-set bench)jev-frontier-bench (4-set average and cost per 1,000) ↗

Plainly: on paid models, Curva costs more per decision than Jev's third-party figure (gpt-4.1-nano $0.078 vs $0.025 per 1,000). Curva is cheaper on free tiers within their daily limits, on repeat decisions (cache) and on questions your rules answer ($0). Curva figures are truncated, not rounded up.

Where Jev is ahead

Jev is one model trained for this job, and it shows on speed and on raw calibration. These are the gaps Curva is working on.

Gaps, Curva runs 2026-10-01
GapCurvaJev
Response time on every model but Groq720 ms to 13,373 ms on PhishNChips; 765 ms to 2,002 ms on AITA (n = 20 to 100)239 ms (PhishNChips); 390 ms (AITA)
Calibration error on BoolQ and multiple-choice sets0.032 to 0.118 (Gemini, Groq; n = 98 to 127)0.024 to 0.038
Phishing calibration error on Groq0.209 after tuning, 0.283 before (n = 100)0.154
OpenBookQA accuracy on Groq85.0%, range 78% to 92% (n = 100)94.2%
AITA accuracyGroq 55.7%, range 46% to 65% (n = 100); OpenAI small models about 41% (early, n = 20 each)75.4%
AITA Brier score (lower is better)Gemini 0.878, Groq 1.042 (n = 100)0.369
Cost per decision on paid models$0.078 (gpt-4.1-nano) to $1.64 (Claude Haiku 4.5) per 1,000$0.025 per 1,000 (third-party, 4-set bench)
Size of the evidence20 to 127 rows per result77 to 2,000 items per published number

How we measured

  • Curva ran with option-order debiasing on, over public dataset rows sampled with a fixed seed. Samples are balanced across labels; accuracy is then weighted back to each label's share of the full public set, which is what a published accuracy measures.
  • The 95% range is on that weighted accuracy. A win needs the whole range above Jev's number.
  • Jev numbers are their published figures, on different samples, mostly measured by third parties. Compare with care.
  • Free models ran 69 to 127 rows per set. Paid models ran 20 per set (early, about ±20 points), so they are not used as evidence here.
  • Calibration after tuning: each half of the rows is tuned on the other half. Reported only from 30 rows up.
  • Curva is closed source and free to use; we publish the datasets, models, sample sizes and dates rather than the harness. Curva docs ↗

Sources for Jev's numbers

Try Curva on your own data

Use the AI API key you already have. Add Curva with a few lines of code. Your AI makes quick decisions you can trust, and Curva itself is free.

pip install curva-ai

Get startedRead the docs ↗

Jev and TypeSafe are trademarks of their respective owners. Tarkova is not affiliated with them. Jev numbers are their published figures.