Curva vs Jev: benchmarks, calibration, speed and cost
Both make an AI pick from your options and tell you how sure it is. Here is how they compare on public test sets, with the model, sample size and date next to every number.

Curva is a free server you run yourself. It gets typed answers with a probability for every option from any LLM you choose, and tunes that confidence from your corrections.
Jev is TypeSafe AI's "System One" model: one trained, closed model behind a hosted API that returns typed answers with probabilities.
With Curva, you use the AI API key you already have (OpenAI, Anthropic, Gemini, Groq, OpenRouter, or a local model), add a few lines of code, and your AI makes quick decisions you can trust. Curva itself is free.
In short: Curva is ahead on PhishNChips accuracy, calibration and speed (n = 100). It matches Jev on five more test sets. Jev is ahead on response time with most models, on calibration error for multiple-choice sets, on AITA, and on cost per decision with paid models.
Benchmarks run 2026-10-01. Page updated .
Where Curva beats Jev
Measured wins only: Curva's whole 95% range sits above Jev's published number, or the benchmark record marks it a win. All on PhishNChips, a public phishing email set, n = 100, run 2026-10-01.
-
80.8%
Jev 62.6%
Accuracy: how often it picks the right answer
-
0.138
Jev 0.154
Calibration error: how far confidence is from reality, lower is better
-
72.8%
Jev 62.6%
Accuracy, on a fast free model
-
178 ms
Jev 239 ms
Typical response time (median)
Also ahead, too small or narrow to lead with: on AITA, Groq's typical response time is 265 ms vs Jev's 390 ms (n = 100), but its AITA accuracy is behind.
Accuracy on every test set
The dot is Curva's accuracy; the bar is its 95% range, where the true score likely sits at this sample size. The outlined mark is Jev's published number.
- Curva, win
- Curva, other results
- 95% range
Jev (published)
-
PhishNChipsGemini flash-lite · n = 100
80.8%Jev 62.6%Curva wins
-
PhishNChipsGroq qwen3.8-27b · n = 100
72.8%Jev 62.6%Curva wins
-
BANKING77Gemini flash-lite · n = 120
79.8%Jev 75.3%matches
-
BANKING77Groq qwen3.8-27b · n = 69
74.7%Jev 75.3%within margin
-
BoolQGemini flash-lite · n = 127
88.9%Jev 89.7%matches
-
OpenBookQAGemini flash-lite · n = 100
92.2%Jev 94.2%matches
-
OpenBookQAGroq qwen3.8-27b · n = 100
85.0%Jev 94.2%Jev ahead
-
CommonsenseQAGroq qwen3.8-27b · n = 100
84.0%Jev 88.1%matches
-
CommonsenseQAGemini flash-lite · n = 100
82.0%Jev 88.1%within margin
-
HellaSwagGroq qwen3.8-27b · n = 100
86.0%Jev 86.1%matches
-
HellaSwagGemini flash-lite · n = 98
78.6%Jev 86.1%within margin
-
AITAGemini flash-lite · n = 100
67.2%Jev 75.4%within margin
-
AITAGroq qwen3.8-27b · n = 100
55.7%Jev 75.4%Jev ahead
Every number, side by side
Calibration error measures how far stated confidence is from how often the answer is right; lower is better. "After tuning" is measured on rows the tuning did not see, which is what you get once your corrections flow in.
| Test set | Curva model | n | Curva accuracy (95% range) | Result | Curva calibration error, before / after tuning | Jev calibration error | Curva typical time (ms) | Jev typical time (ms) | |
|---|---|---|---|---|---|---|---|---|---|
| PhishNChips | Gemini flash-lite | 100 | 80.8% (73% to 89%) | 62.6% | Curva wins | 0.138 / 0.138 | 0.154 | 1,146 | 239 |
| PhishNChips | Groq qwen3.8-27b | 100 | 72.8% (64% to 82%) | 62.6% | Curva wins | 0.283 / 0.209 | 0.154 | 178 | 239 |
| BANKING77 | Gemini flash-lite | 120 | 79.8% (73% to 87%) | 75.3% | matches | 0.097 / 0.097 | not published | 1,362 | not published |
| BANKING77 | Groq qwen3.8-27b | 69 | 74.7% (64% to 85%) | 75.3% | within margin | 0.252 / 0.157 | not published | 351 | not published |
| BoolQ | Gemini flash-lite | 127 | 88.9% (83% to 94%) | 89.7% | matches | 0.081 / 0.081 | 0.038 | 1,071 | not published |
| OpenBookQA | Gemini flash-lite | 100 | 92.2% (87% to 97%) | 94.2% | matches | 0.056 / 0.056 | 0.024 | 1,013 | not published |
| OpenBookQA | Groq qwen3.8-27b | 100 | 85.0% (78% to 92%) | 94.2% | Jev ahead | 0.051 / 0.051 | 0.024 | 207 | not published |
| CommonsenseQA | Groq qwen3.8-27b | 100 | 84.0% (77% to 91%) | 88.1% | matches | 0.066 / 0.066 | 0.032 | 230 | not published |
| CommonsenseQA | Gemini flash-lite | 100 | 82.0% (74% to 90%) | 88.1% | within margin | 0.118 / 0.118 | 0.032 | 1,263 | not published |
| HellaSwag | Groq qwen3.8-27b | 100 | 86.0% (79% to 93%) | 86.1% | matches | 0.085 / 0.085 | 0.029 | 209 | not published |
| HellaSwag | Gemini flash-lite | 98 | 78.6% (70% to 87%) | 86.1% | within margin | 0.032 / 0.032 | 0.029 | 1,173 | not published |
| AITA | Gemini flash-lite | 100 | 67.2% (58% to 76%) | 75.4% | within margin | 0.404 / 0.149 | not published | 1,318 | 390 |
| AITA | Groq qwen3.8-27b | 100 | 55.7% (46% to 65%) | 75.4% | Jev ahead | 0.475 / 0.221 | not published | 265 | 390 |
On PhishNChips, Gemini's calibration error (0.138) and Groq's response time (178 ms) are wins; Groq's calibration error (0.209 after tuning) is behind Jev's 0.154. Jev publishes no calibration error for AITA; its AITA Brier score is 0.369 against Curva's 0.878 (Gemini) and 1.042 (Groq). Paid models (gpt-4.1-nano, gpt-4o-mini, gpt-4.1-mini, Claude Haiku 4.5) ran only 20 rows per set, too early to compare, so they are left out of this table.
Features: Curva vs Jev
| Area | ||
|---|---|---|
| The model | Any OpenAI-compatible model: OpenAI, Anthropic, Gemini, Groq, DeepSeek, Mistral, OpenRouter, or local Ollama, vLLM, LM Studio, llama.cpp. Mix them in one request. | One closed model, Jev. Architecture, weights and paper undisclosed. TypeSafe launch post ↗ |
| Where it runs | Your servers: one binary with SQLite, or Docker. Your data stays with you. | TypeSafe's hosted API, served from the US West Coast. TypeSafe launch post ↗ |
| Price | Free to use, also commercially. You pay only your model provider, and free models work. | $0.042 per 1M input tokens, output free (self-reported). TypeSafe launch post ↗ |
| Answer types | Choice (2 to 255 options), Score (2 to 20 levels), yes/no, multi-select, text, number, integer. | Choice (up to 255 options), Score (2 to 10 levels), yes/no probability. Jev docs ↗ |
| Images | Up to 8 per decision, with vision models. | Text only. Jev docs ↗ |
| When unsure | A "none of these" option by default. Below your threshold it says it is not sure, so a person decides. Or a short list that holds the right answer 95% of the time. | No "none of these" option. In one test, 7 of 53 answers reached 0.9 confidence. AgentConn review ↗ |
| Confidence on your data | Tuned from your corrections, per question. Kept only when it helps on rows it did not learn from. | Fixed at training (self-reported). In one test a fair coin came out 0.92 heads. Alex Molas: "Jev can't be calibrated" ↗ |
| Linked questions | One question can depend on another (depends_on, when), plus rules and per-question reasoning. | Not in Jev's docs. Jev docs ↗ |
| Several models | Fallback, council, cascade and race across models. | One model. TypeSafe launch post ↗ |
| Audit | An audit log for every decision, and which input fields moved the answer. | No reason returned with a decision. AgentConn review ↗ |
| Versions | Pinned configs. You move curva-latest yourself. | jev-latest moves to new releases; a pinned version is available. Jev docs ↗ |
| Typical speed | 178 ms on Groq (PhishNChips, n = 100). 720 ms to 13,373 ms on other models. Repeat decisions come from a cache at $0. | 70 to 500 ms (self-reported). 239 ms measured on PhishNChips. jev-phishing-bench (PhishNChips, 2,000 emails) ↗ |
| Input size | Up to 150,000 characters of input. The total limit is your model's. | 64k tokens; input plus the longest question within 32k. Jev docs ↗ |
| SDKs | Python (sync and async), TypeScript, n8n node, MCP server, OpenAPI. | Python and TypeScript. Jev docs ↗ |
Cost per 1,000 decisions
Curva is free to use and runs on your servers, so there is no per-call fee to Tarkova. You pay your model provider. Each decision is two model calls by default (the options are asked in both orders).
| Model | Per 1,000 decisions | Note |
|---|---|---|
| Gemini flash-lite | $0 | free tier, about 250 decisions a day |
| Groq qwen3.8-27b | $0 | free tier, about 180 decisions a day |
| gpt-4.1-nano | $0.078 | $0.055 on PhishNChips |
| gpt-4o-mini | $0.118 | $0.083 on PhishNChips |
| gpt-4.1-mini | $0.314 | $0.223 on PhishNChips |
| Claude Haiku 4.5 | $1.64 | $1.29 on PhishNChips |
| Any model, repeat decision | $0 | served from the cache |
| Figure | Value | Source |
|---|---|---|
| List price | $0.042 per 1M input tokens, output free (self-reported) | TypeSafe launch post ↗ |
| Per 1,000 decisions | $0.025 (third-party, 4-set bench) | jev-frontier-bench (4-set average and cost per 1,000) ↗ |
Plainly: on paid models, Curva costs more per decision than Jev's third-party figure (gpt-4.1-nano $0.078 vs $0.025 per 1,000). Curva is cheaper on free tiers within their daily limits, on repeat decisions (cache) and on questions your rules answer ($0). Curva figures are truncated, not rounded up.
Where Jev is ahead
Jev is one model trained for this job, and it shows on speed and on raw calibration. These are the gaps Curva is working on.
| Gap | Curva | |
|---|---|---|
| Response time on every model but Groq | 720 ms to 13,373 ms on PhishNChips; 765 ms to 2,002 ms on AITA (n = 20 to 100) | 239 ms (PhishNChips); 390 ms (AITA) |
| Calibration error on BoolQ and multiple-choice sets | 0.032 to 0.118 (Gemini, Groq; n = 98 to 127) | 0.024 to 0.038 |
| Phishing calibration error on Groq | 0.209 after tuning, 0.283 before (n = 100) | 0.154 |
| OpenBookQA accuracy on Groq | 85.0%, range 78% to 92% (n = 100) | 94.2% |
| AITA accuracy | Groq 55.7%, range 46% to 65% (n = 100); OpenAI small models about 41% (early, n = 20 each) | 75.4% |
| AITA Brier score (lower is better) | Gemini 0.878, Groq 1.042 (n = 100) | 0.369 |
| Cost per decision on paid models | $0.078 (gpt-4.1-nano) to $1.64 (Claude Haiku 4.5) per 1,000 | $0.025 per 1,000 (third-party, 4-set bench) |
| Size of the evidence | 20 to 127 rows per result | 77 to 2,000 items per published number |
How we measured
- Curva ran with option-order debiasing on, over public dataset rows sampled with a fixed seed. Samples are balanced across labels; accuracy is then weighted back to each label's share of the full public set, which is what a published accuracy measures.
- The 95% range is on that weighted accuracy. A win needs the whole range above Jev's number.
- Jev numbers are their published figures, on different samples, mostly measured by third parties. Compare with care.
- Free models ran 69 to 127 rows per set. Paid models ran 20 per set (early, about ±20 points), so they are not used as evidence here.
- Calibration after tuning: each half of the rows is tuned on the other half. Reported only from 30 rows up.
- Curva is closed source and free to use; we publish the datasets, models, sample sizes and dates rather than the harness. Curva docs ↗
Sources for Jev's numbers
- TypeSafe launch post ↗
- Jev docs ↗
- TypeSafe evals ↗
- jev-phishing-bench (PhishNChips, 2,000 emails) ↗
- llmevals: Jev on BANKING77 (77 items) ↗
- jev-ood-calibration (OpenBookQA, CommonsenseQA, HellaSwag) ↗
- jev-aita (AITA, 770 posts) ↗
- jev-frontier-bench (4-set average and cost per 1,000) ↗
- Alex Molas: "Jev can't be calibrated" ↗
- AgentConn review ↗
- dev.to guide to Jev ↗
- BoolQ: Nimble PUBLIC_BENCHMARKS boolq subset (89.7%, calibration error 0.038)
Try Curva on your own data
Use the AI API key you already have. Add Curva with a few lines of code. Your AI makes quick decisions you can trust, and Curva itself is free.
pip install curva-ai
Jev and TypeSafe are trademarks of their respective owners. Tarkova is not affiliated with them. Jev numbers are their published figures.