curva.local() vs Curva(): running the server from Python
Three ways to get a Curva client in Python: module-level decide, curva.local() with its own server, or Curva() for a shared one. What each does.
Curva runs a local LLM classification server from Python in three ways. The module-level curva.decide starts a private server on its first call and stops it when Python exits. curva.local() does the same but hands you the lifecycle: a client bound to a server on a free 127.0.0.1 port, its own database file, and a with block to stop it. Curva(), given a base URL, starts nothing and talks to a shared server you run with curva serve or Docker. All three run the same curva binary, which ships inside the Python wheel. This guide shows what each one does with your environment, your database and your keys, and when to move from one to the next.
A local LLM classification server from Python: three entry points
flowchart LR A["curva.decide(...)"] --> P1["private server, started on first call"] B["curva.local()"] --> P2["private server on a free 127.0.0.1 port"] C["Curva(base_url)"] --> S["shared server: curva serve or Docker"] P2 --> DB["~/.curva/curva.db or db="] S --> DB2["the server's own --db file"]
| Entry point | Starts a server | Lifecycle | Fits |
|---|---|---|---|
curva.decide | yes, on the first call | stops when Python exits | scripts, notebooks, the three-line demo |
curva.local() | yes | Python exit, close(), or the end of a with block | tests, local tools, anything that wants its own database |
Curva() with a base URL | no | yours | services, teams, anything with more than one process |
curva.decide: a private server on the first call, stopped when Python exits
import curva
d = curva.decide("I was charged twice, please refund me",
{"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total) # billing True NoneThere is nothing to set up. The first curva.decide starts a private server for this Python process, reuses it for every later call, and stops it when Python exits. curva.feedback and curva.calibration at module level use the same server.
One switch changes that: with CURVA_BASE_URL set, the module-level functions talk to that server instead of starting their own. So the same script can run against a private server on a laptop and a shared one in production, with only an environment variable between them.
curva.local(): a free 127.0.0.1 port, startup_timeout=15.0 and the with block
curva.local(model=None, db=None, *, startup_timeout=15.0, **client_options)curva.local() starts curva serve on a free 127.0.0.1 port and returns a client bound to it. Because it listens only on localhost, nothing outside your machine can reach it, and it needs no API key.
modelsets the server's default model.dbpicks the SQLite file for decisions, feedback and calibrators.startup_timeoutis how long to wait for the server to come up, 15 seconds by default.client_optionsgo to the client:timeout,max_retries,api_key.
The server stops when Python exits, when you call close(), or at the end of a with block:
with curva.local() as client:
...The full first decision from the docs uses it this way:
import curva
from curva import Choice, Score, Noul
client = curva.local()
d = client.decide(
state={"ticket": "I was charged twice for order A-104. Please refund the duplicate!"},
questions={
"team": Choice("Which team should handle this?",
{"billing": "payments, refunds", "technical": "bugs", "sales": "pricing"}),
"frustration": Score("How frustrated is the customer?", ["calm", "annoyed", "angry"]),
"refund": Noul("The customer explicitly asks for a refund"),
},
)
print(d["team"].choice, d["team"].confidence) # billing 0.9999
print(d["frustration"].score) # 0.65
print(d["refund"].noul) # 0.999
print(d.mode, d.latency_ms, d.cost_usd, d.cached)Where calibration lives: ~/.curva/curva.db and the db= argument
Calibration is only useful if it survives. curva.local() keeps its database at ~/.curva/curva.db by default, so calibration learned in one run is there in the next. Feedback you send today fits a calibrator that tomorrow's script uses.
Pass db= to keep separate databases: one per experiment, a throwaway file in tests, or a fixed path in a container. Calibrators are also kept per project inside one database, so two use cases can share a file without mixing labels.
The server inherits your environment: OPENROUTER_API_KEY, CURVA_MODEL, CURVA_RPM
The private server is a child of your Python process, and it inherits the process's environment. So set provider variables before you start it:
export OPENROUTER_API_KEY=sk-or-v1-... # any OpenRouter key; free models workWithout OPENROUTER_API_KEY, the first key set among OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, GROQ_API_KEY and a few others picks a small, fast model of that provider. CURVA_MODEL chooses one yourself, for example CURVA_MODEL=@ollama/qwen3:4b for a fully local setup. Server settings such as CURVA_RPM pass through the same way. With no key at all, curva.decide raises an error naming the variables to set.
Moving to a shared server with CURVA_BASE_URL and CURVA_API_KEY
A private server per process is right for a script. For a service, it means one database, one decision cache and one set of calibrators per process, and none of them shared. Move to one shared server:
curva serve --addr 127.0.0.1:7777 --db curva.db
curva keys create --name my-servicefrom curva import Curva
client = Curva("http://your-server:7777") # or set CURVA_BASE_URLCurva() reads CURVA_BASE_URL, default http://localhost:7777, and CURVA_API_KEY when they aren't passed. Once the server has API keys, every request needs one. A shared server keeps one audit log, one set of calibrators per project, and one decision cache, so a repeat from any client comes back in about a millisecond. For production, the Docker image runs the same binary with its data in a volume.
Three mistakes to avoid
**A private server per request.** curva.local() starts a whole server. Calling it inside a web request handler starts one per request, each with its own cache and its own view of the database. Create one client when your process starts and reuse it, or point every process at one shared server with Curva().
**Tests that write to your real calibration.** By default curva.local() writes to ~/.curva/curva.db. A test suite that sends feedback would then fit calibrators your scripts later use. Give tests their own file with db=, and their own project.
**Several processes, one file, no server.** Two scripts that each start curva.local() on the same database file each run their own server against it. If several processes need the same calibrators and cache, that is the moment to run one curva serve and connect with Curva(). The server has one writer, one cache and one audit log, which is what a team needs.
Which wheels ship the curva binary: Linux, macOS and Windows
curva.local() and module-level curva.decide need the curva binary, and pip install curva-ai provides it. Platform wheels for Linux (x86_64, aarch64), macOS (Apple Silicon, Intel) and Windows include the server binary, so you don't need Rust. The SDK itself uses only the standard library, on Python 3.9 or newer. curva --version should print curva 0.1.0.
Next steps
The docs cover the clients in the Python SDK reference and getting started. Start with the Python LLM classification tutorial, learn the shorthand in Python LLM shorthand questions, and tune the client in LLM client retries and timeouts in Python. When you outgrow a private server, read self-host an LLM decision API.