Switch from Jev to Curva: a step-by-step guide
Move a Jev integration to Curva: /v1/systemone, the criteria shape, two environment variables, a shadow test, and the differences to expect.
To switch from Jev to Curva, run a Curva server, create an API key, and point your existing client at it: Curva serves POST /v1/systemone, which accepts Jev's criteria request shape, so a client built for Jev keeps working after you change two environment variables. Then replay logged traffic through Curva with a shadow test, read where the two disagree, and move traffic once the numbers say so. This guide walks the steps in order, says what was verified and what wasn't, and lists the differences to expect before you cut over.
Jev by TypeSafe AI is a hosted "System One" model: unstructured state in, typed probabilistic decisions out. Curva gives the same kind of output from any LLM you choose, on your own servers. It is free to use under the Curva Free License; you pay only your model provider.
Run Curva and create a key
curva serve --addr 127.0.0.1:7777
curva keys create --name my-app # prints curva_… once
export TYPESAFE_BASE_URL=http://127.0.0.1:7777
export TYPESAFE_API_KEY=curva_...pip install curva-ai installs the curva binary. The server needs a model provider key, such as OPENROUTER_API_KEY, or a local model. Curva doesn't ship a model of its own, so pick one before you compare and write it down: every number you measure belongs to that model.
Two details about keys. Once any key exists, every route except /health needs one, and Jev's SDK always sends one. And a server started without keys stays open until it is restarted, so restart it after creating the first key. For production, the Docker image ghcr.io/itsmohitrohilla/curva runs the same binary.
Point the client at /v1/systemone
You have two ways to switch from Jev.
**Keep Jev's Python SDK.** Set TYPESAFE_BASE_URL to your Curva server and TYPESAFE_API_KEY to a Curva key, as above. Nothing else changes. This was verified with typesafe-sdk 0.7.2 against curva serve: system_one, all three primitives with criteria, dict questions, extra_body, response.choices/scores/nouls, usage, models.list(), and 401 and 422 errors all behaved. Jev's TypeScript SDK uses the same wire contract and should work the same way, but it was not run. Test it before you rely on it.
**Move to Curva's SDK.** Its Python client accepts the same style, so most code survives an import change:
from curva import Curva as TypeSafeClient, CurvaError as TypeSafeError, Choice, Score, Noul
with TypeSafeClient(base_url="http://127.0.0.1:7777", api_key="curva_...") as client:
result = client.system_one(
state={"ticket": "I was charged twice for order A-104"},
questions={
"department": Choice(instructions="Which team",
criteria={"billing": "payments, refunds", "technical": "bugs"}),
"frustration": Score(instructions="How frustrated", criteria=["Calm", "Frustrated", "Angry"]),
"refund_requested": Noul(instructions="Customer explicitly asks for a refund"),
},
)
print(result.choices["department"].choice)Pass base_url and api_key by keyword. Also drop any model= argument that names jev-latest: through Curva's own SDK a model name is sent to the provider as is, and your provider doesn't know that name. On /v1/systemone, a name without a provider prefix means the server's default model.
Questions in the criteria shape
If you call the HTTP API yourself, the body stays in the criteria shape:
{
"state": {"ticket": "I was charged twice for order A-104"},
"model": "curva-latest",
"questions": {
"team": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "payments, refunds", "technical": "bugs", "sales": null}},
"anger": {"type": "score", "instructions": "How frustrated?",
"criteria": ["calm", "annoyed", "angry"]},
"refund": {"type": "noul", "instructions": "Asks for a refund",
"criteria": {"true": "explicitly asks for money back", "false": "anything else"}}
}
}How Curva reads each part:
- **Choice
criteria** maps a key to a description: text,nullor JSON. On this route there is nonone_of_theseescape option unless the question sets"escape": true, so answers stay inside your criteria, as your client expects. - **Score
criteria** lists level descriptions, lowest first. Curva takes 2 to 20 levels; Jev takes 2 to 10. - **Noul
criteria** is optional:{"true": ..., "false": ...}. - **
instructions** may be omitted, or be JSON. - **
{"type": "string"}** withoutenumis answered as a text question. - **
model** set tocurva-latestor a pinnedcurva-x.y.zselects a config. Anything with a provider prefix is a model id or a multi-model plan.
The response is Curva's normal response plus usage and a type on every answer. A Score's probabilities and legend are keyed "0", "1" and so on, as Jev's SDK expects.
Shadow-test on logged traffic
Don't move traffic on faith. Log a day of Jev's decisions, one per line, with its answers as labels in feedback format (option key, level index, or true and false):
{"state": {"ticket": "I was charged twice"}, "labels": {"team": "billing", "refund_requested": true}}Then replay it:
curva shadow traffic.jsonl -q questions.json -o shadow.jsonlquestions.json holds the same questions in Curva's /v1/decide format, with options and levels instead of criteria. The Markdown report shows agreement per question, agreement by Curva's confidence band, and the share of traffic you could hand to Curva at or above each band. The most confident disagreements come first. In each one, one of the two systems is clearly wrong, so read them. Agreement is not accuracy: label a sample of the disagreements yourself. The run is resumable, so a quota error costs nothing but time.
flowchart LR
A["run curva serve, create a key"] --> B["point the client at /v1/systemone"]
B --> C["curva shadow on logged Jev traffic"]
C --> D{"agreement high enough at a confidence band?"}
D -- yes --> E["move that band of traffic"]
D -- no --> F["read disagreements, label a sample"]
F --> CDifferences: headers, error codes, usage
Know these before you cut over:
| Area | Jev | Curva |
|---|---|---|
| Request id header | x-typesafe-request-id | Not sent; the decision id is id in the body, and response.request_id raises in Jev's SDK |
| Malformed JSON | 422 | 400 |
| Upstream model failure | 529 | 502 |
usage.output_tokens | tracked | always 0; input tokens are summed over calls |
GET /v1/models | model list | Curva's pinned configs, with an empty release_date |
| Latency | 70 to 500 ms, self-reported | about 0.3 to 3 s on remote models; about 1 ms on cache hits |
| Calls per decision | one | two by default, for order debiasing |
Latency is the big one. If you need sub-second answers, test a fast provider such as Groq, a local model, rules for the easy cases, and debias: "auto", which drops to one call once a question has shown no position bias.
Turn on what Jev didn't have
Once traffic flows, add the parts Curva adds:
- **Feedback.** Store the decision
idand send the true answer withPOST /v1/feedback. After 30 labels per question, Curva calibrates when that helps. - **Abstain.** Set
min_confidencewhere a wrong answer is expensive, and sendabstain: trueanswers to a person. - **Coverage sets.** Set
coveragefor a set of options with a guaranteed hit rate, once a question has 30 labels. - **Pinned configs.** Pin
curva-1.1.0so answers don't shift when you upgrade. - **Audit log.** Every decision is logged with model, config and answers, and a salted hash of the state, never the state.
The checklist to switch from Jev
- Each question mapped to Choice, Score, Noul or Multi.
- Curva running, API key created, server restarted after the first key.
- Base URL changed to the Curva server.
modelremoved or set to a Curva config.- Shadow run on logged traffic, disagreements read and a sample labeled.
- Decision ids stored, feedback flowing.
min_confidenceset where mistakes cost money.
Next steps
The docs have the switching guide and shadow mode. For measured accuracy, speed and cost side by side, read a self-hosted Jev alternative or the Curva vs Jev page. For the shadow test in depth, see shadow-test an LLM classifier, and for what feedback buys you, LLM classification confidence scores you can act on.