Ollama classification with probabilities, fully local

Run typed LLM classification on your own machine with Ollama: no key, no network hop, privacy strict allowed, and a cost of $0 per call.

Ollama classification with Curva gives you typed answers from a model on your own machine: one of your labels per question, with a probability for every option, and no API key, no network hop and no per-call price. You pull a small model with ollama pull qwen3:4b, set CURVA_MODEL=@ollama/qwen3:4b, and call curva.decide. Because the model's URL is on your machine, Curva also accepts privacy: "strict", which turns "nothing leaves the machine" from a habit into a rule. This guide covers the setup, why strict works locally and fails on a remote provider, how to test a small model before you trust it, and how to move Ollama to another port or host.

Some data shouldn't leave the building: customer messages under contract, internal tickets, anything your legal team has opinions about. Some teams just don't want a per-call bill or a rate limit. For both, the answer is a model you run yourself. Curva is free to use under the Curva Free License, and a local model has no per-call price, so the running cost is your hardware.

ollama pull qwen3:4b and CURVA_MODEL

Ollama serves an OpenAI-compatible API on http://127.0.0.1:11434/v1, which is the default URL of Curva's built-in ollama provider. Two commands set it up:

bash
ollama pull qwen3:4b
export CURVA_MODEL=@ollama/qwen3:4b

Then install the Python package, which includes the Curva server: pip install curva-ai. That is the whole configuration. Model ids that start with @ollama/ go to the built-in Ollama provider, which needs no key, so you don't set OPENROUTER_API_KEY or any other provider key. qwen3:4b is the docs' example; any instruct model you have pulled works the same way.

Your first Ollama classification with privacy strict

python
import curva
d = curva.decide("The app crashes on launch", {"team": ["billing", "technical"]}, privacy="strict")

curva.decide starts a private Curva server for this Python process on a free localhost port, reuses it for later calls, and stops it when Python exits. The model call goes to Ollama on localhost. The answer is typed: d.team is billing or technical, and d["team"] holds the confidence and a probability for each option. d["team"] can also be none_of_these, the escape option every Choice gets, when the text fits neither team.

The local database lives at ~/.curva/curva.db, so calibration learned in one run is there in the next. On disk, the audit log keeps only a salted hash of each state, never the state itself.

Why strict works locally and fails on a remote provider

privacy: "strict" routes the model call only to providers that neither store nor train on prompts. How Curva enforces it depends on the provider:

  • On OpenRouter, strict sets data_collection: deny.
  • On a named provider such as @ollama, @openai or @groq, strict is accepted only when the provider's URL is on this machine. Prompts never leave it.
  • Otherwise the request gets 422.

That last line is the point. If someone later points the model at a hosted API by mistake, a strict request fails loudly instead of quietly sending data out. Strict also returns extracted values (Text, Number, Integer) without storing them; the audit log shows them as redacted.

flowchart LR
  A["curva.decide(..., privacy=strict)"] --> B{"provider URL on this machine?"}
  B -- "yes: @ollama on 127.0.0.1" --> C["call Ollama locally, typed answer"]
  B -- "no: a hosted API" --> D["422, nothing sent"]
Figure 1. Ollama classification with privacy strict Strict passes for a provider URL on this machine and is refused with 422 for a remote one.

Local models also change the bookkeeping. cost_usd is always 0 on your own machine. If you want decisions to carry your GPU cost, give a named provider a price with CURVA_PROVIDER_<NAME>_PRICE (input and output dollars per million tokens). And a provider URL on this machine has no rate limit by default.

Speed and accuracy: test a small local model first

A 4B model is not a frontier model. Measure it on your data before you route anything on it.

**Check the probabilities.** Curva reads probabilities in logprobs mode (one label token per question, probabilities from the token log-probabilities) or verbal mode (JSON limited to your labels, with a probability for each). The default, auto, tries logprobs first, falls back to verbal, and remembers the result per model id. Whether a local model returns logprobs varies, so check it:

bash
curva spike @ollama/qwen3:4b

curva spike reports whether a model gives usable label probabilities. Either mode works. The response's mode field tells you which one answered.

**Measure accuracy.** Run curva bench on a labeled set with --model @ollama/qwen3:4b to see accuracy, calibration error, latency and cost, or replay your own logged traffic with curva shadow.

**Help the small model.** A few settings matter more locally:

  • Order debiasing is on by default. Each question is asked with the options in both orders and averaged, which cancels a small model's lean towards one position. It costs two local calls; debias="auto" drops to one once a question has shown no position bias.
  • Few-shot examples, up to 10 per question, fix the cases it keeps getting wrong:
python
from curva import Noul, Choice

urgent = Noul("Is this urgent?", examples=[
    ({"ticket": "The production server is down"}, True),
    ({"ticket": "Typo on the pricing page"}, False),
])

team = Choice("Which team?", ["billing", "technical"], examples=[
    ({"ticket": "Refund my last invoice"}, "billing"),
])
  • Feedback calibrates it. Send true answers with client.feedback(d.id, "team", label). After 30 labels per question, Curva fits a calibrator and keeps it when it improves held-out accuracy of the probabilities. The calibrator lives in the local SQLite file like everything else.
  • min_confidence on questions that route work sends unsure answers to a person.

Moving Ollama to another port or host

If Ollama listens somewhere else, move the built-in provider with one variable:

bash
export CURVA_PROVIDER_OLLAMA_URL=http://127.0.0.1:9999/v1 # moves the built-in @ollama

The same works for a GPU box on your network: CURVA_PROVIDER_GPU_URL=http://gpu-box:8000/v1 adds a provider named gpu, and model ids become @gpu/<model>. Remember that privacy: strict passes only when that URL is on this machine, so a model on another host is a remote provider as far as strict is concerned.

For more than one script, run one Curva server and point clients at it:

bash
curva serve --model @ollama/qwen3:4b --db curva.db

Without API keys, curva serve listens only on localhost. To serve other machines, create a key with curva keys create, restart the server, and put a reverse proxy with TLS in front.

The same pattern works for the other local servers Curva knows:

ProviderDefault URLModel id
ollamahttp://127.0.0.1:11434/v1@ollama/qwen3:4b
lmstudiohttp://127.0.0.1:1234/v1@lmstudio/qwen3-4b
vllmhttp://127.0.0.1:8000/v1@vllm/Qwen/Qwen3-4B
llamacpphttp://127.0.0.1:8080/v1@llamacpp/model

When to add a hosted model

Some teams keep everything local. Others want a stronger model for the hard minority. A cascade asks the local model first and sends only the questions it is unsure about, below escalate_below (default 0.8), to the next model. A council can mix a local and a hosted model too. Be clear about what that means: those states leave the machine, and a strict request will refuse a named hosted provider. Decide per project which data may go out.

Next steps

The docs cover this in model providers and getting started. To compare a local model with hosted ones, read choose a model for LLM classification. For the free hosted route, see free LLM APIs for classification, and to check logprobs support first, check LLM logprobs support.