Ollama classification with probabilities, fully local
Run typed LLM classification on your own machine with Ollama: no key, no network hop, privacy strict allowed, and a cost of $0 per call.
Ollama classification with Curva gives you typed answers from a model on your own machine: one of your labels per question, with a probability for every option, and no API key, no network hop and no per-call price. You pull a small model with ollama pull qwen3:4b, set CURVA_MODEL=@ollama/qwen3:4b, and call curva.decide. Because the model's URL is on your machine, Curva also accepts privacy: "strict", which turns "nothing leaves the machine" from a habit into a rule. This guide covers the setup, why strict works locally and fails on a remote provider, how to test a small model before you trust it, and how to move Ollama to another port or host.
Some data shouldn't leave the building: customer messages under contract, internal tickets, anything your legal team has opinions about. Some teams just don't want a per-call bill or a rate limit. For both, the answer is a model you run yourself. Curva is free to use under the Curva Free License, and a local model has no per-call price, so the running cost is your hardware.
ollama pull qwen3:4b and CURVA_MODEL
Ollama serves an OpenAI-compatible API on http://127.0.0.1:11434/v1, which is the default URL of Curva's built-in ollama provider. Two commands set it up:
ollama pull qwen3:4b
export CURVA_MODEL=@ollama/qwen3:4bThen install the Python package, which includes the Curva server: pip install curva-ai. That is the whole configuration. Model ids that start with @ollama/ go to the built-in Ollama provider, which needs no key, so you don't set OPENROUTER_API_KEY or any other provider key. qwen3:4b is the docs' example; any instruct model you have pulled works the same way.
Your first Ollama classification with privacy strict
import curva
d = curva.decide("The app crashes on launch", {"team": ["billing", "technical"]}, privacy="strict")curva.decide starts a private Curva server for this Python process on a free localhost port, reuses it for later calls, and stops it when Python exits. The model call goes to Ollama on localhost. The answer is typed: d.team is billing or technical, and d["team"] holds the confidence and a probability for each option. d["team"] can also be none_of_these, the escape option every Choice gets, when the text fits neither team.
The local database lives at ~/.curva/curva.db, so calibration learned in one run is there in the next. On disk, the audit log keeps only a salted hash of each state, never the state itself.
Why strict works locally and fails on a remote provider
privacy: "strict" routes the model call only to providers that neither store nor train on prompts. How Curva enforces it depends on the provider:
- On OpenRouter, strict sets
data_collection: deny. - On a named provider such as
@ollama,@openaior@groq, strict is accepted only when the provider's URL is on this machine. Prompts never leave it. - Otherwise the request gets 422.
That last line is the point. If someone later points the model at a hosted API by mistake, a strict request fails loudly instead of quietly sending data out. Strict also returns extracted values (Text, Number, Integer) without storing them; the audit log shows them as redacted.
flowchart LR
A["curva.decide(..., privacy=strict)"] --> B{"provider URL on this machine?"}
B -- "yes: @ollama on 127.0.0.1" --> C["call Ollama locally, typed answer"]
B -- "no: a hosted API" --> D["422, nothing sent"]Local models also change the bookkeeping. cost_usd is always 0 on your own machine. If you want decisions to carry your GPU cost, give a named provider a price with CURVA_PROVIDER_<NAME>_PRICE (input and output dollars per million tokens). And a provider URL on this machine has no rate limit by default.
Speed and accuracy: test a small local model first
A 4B model is not a frontier model. Measure it on your data before you route anything on it.
**Check the probabilities.** Curva reads probabilities in logprobs mode (one label token per question, probabilities from the token log-probabilities) or verbal mode (JSON limited to your labels, with a probability for each). The default, auto, tries logprobs first, falls back to verbal, and remembers the result per model id. Whether a local model returns logprobs varies, so check it:
curva spike @ollama/qwen3:4bcurva spike reports whether a model gives usable label probabilities. Either mode works. The response's mode field tells you which one answered.
**Measure accuracy.** Run curva bench on a labeled set with --model @ollama/qwen3:4b to see accuracy, calibration error, latency and cost, or replay your own logged traffic with curva shadow.
**Help the small model.** A few settings matter more locally:
- Order debiasing is on by default. Each question is asked with the options in both orders and averaged, which cancels a small model's lean towards one position. It costs two local calls;
debias="auto"drops to one once a question has shown no position bias. - Few-shot examples, up to 10 per question, fix the cases it keeps getting wrong:
from curva import Noul, Choice
urgent = Noul("Is this urgent?", examples=[
({"ticket": "The production server is down"}, True),
({"ticket": "Typo on the pricing page"}, False),
])
team = Choice("Which team?", ["billing", "technical"], examples=[
({"ticket": "Refund my last invoice"}, "billing"),
])- Feedback calibrates it. Send true answers with
client.feedback(d.id, "team", label). After 30 labels per question, Curva fits a calibrator and keeps it when it improves held-out accuracy of the probabilities. The calibrator lives in the local SQLite file like everything else. min_confidenceon questions that route work sends unsure answers to a person.
Moving Ollama to another port or host
If Ollama listens somewhere else, move the built-in provider with one variable:
export CURVA_PROVIDER_OLLAMA_URL=http://127.0.0.1:9999/v1 # moves the built-in @ollamaThe same works for a GPU box on your network: CURVA_PROVIDER_GPU_URL=http://gpu-box:8000/v1 adds a provider named gpu, and model ids become @gpu/<model>. Remember that privacy: strict passes only when that URL is on this machine, so a model on another host is a remote provider as far as strict is concerned.
For more than one script, run one Curva server and point clients at it:
curva serve --model @ollama/qwen3:4b --db curva.dbWithout API keys, curva serve listens only on localhost. To serve other machines, create a key with curva keys create, restart the server, and put a reverse proxy with TLS in front.
The same pattern works for the other local servers Curva knows:
| Provider | Default URL | Model id |
|---|---|---|
| ollama | http://127.0.0.1:11434/v1 | @ollama/qwen3:4b |
| lmstudio | http://127.0.0.1:1234/v1 | @lmstudio/qwen3-4b |
| vllm | http://127.0.0.1:8000/v1 | @vllm/Qwen/Qwen3-4B |
| llamacpp | http://127.0.0.1:8080/v1 | @llamacpp/model |
When to add a hosted model
Some teams keep everything local. Others want a stronger model for the hard minority. A cascade asks the local model first and sends only the questions it is unsure about, below escalate_below (default 0.8), to the next model. A council can mix a local and a hosted model too. Be clear about what that means: those states leave the machine, and a strict request will refuse a named hosted provider. Decide per project which data may go out.
Next steps
The docs cover this in model providers and getting started. To compare a local model with hosted ones, read choose a model for LLM classification. For the free hosted route, see free LLM APIs for classification, and to check logprobs support first, check LLM logprobs support.