Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSHow to word LLM classification questions: recipe lessons
Wording lessons from Curva's shipped recipes: judge by need, define the no case, carve out false positives, and freeze wording before collecting labels.
More articles
Page 1 of 18-
Classify support tickets that include screenshots
Send a ticket's screenshot to a vision model next to its text: is an error visible, which part of the app. Images inline, up to 8 per request.
-
Reliability diagrams: plot LLM confidence vs accuracy
Read a reliability diagram for an LLM classifier from Curva's reliability bins or dashboard, and spot overconfident, underconfident and leaning models.
-
Detect refund requests, not the word refund
A yes/no question that separates "I was charged twice" from "please refund me": the recipe's wording, the probability it returns and how to calibrate it.
-
Plain vs full LLM answers in Python: the hidden 0.5
d.refund is True whenever P(yes) is at least 0.5. When plain answers are fine, when to read the full answer, and the id and model name clash.
-
Python LLM API error handling with CurvaError
Map each CurvaError subclass to the fix: 422 names the bad question, 429 carries retry_after, 502 splits model_error from model_unavailable.
-
Product catalog tagging with multi-label LLM questions
Tag products with a Multi question: an independent probability per tag, a threshold you tune, and up to 20 tags per question across a whole catalog.
-
Pin an LLM prompt version so answers don't shift
A pinned config freezes the prompt template, mode, debias and model so upgrades don't move your answers. The pins, curva-latest, and when to recalibrate.
-
Ollama classification with probabilities, fully local
Run typed LLM classification on your own machine with Ollama: no key, no network hop, privacy strict allowed, and a cost of $0 per call.
-
None of the above: an escape option for LLM classifiers
A classifier forced to pick a label will, confidently. How a none_of_these option works, why it is on by default, and how to route it to a person.
-
Logprobs vs verbal confidence: which LLM number to trust
Token logprobs or a stated probability per label? How each mode works, which one auto picks, and the probe results that changed Curva's default model.
-
LLM yes or no questions with a real probability of yes
A Noul question returns one number, P(yes), asked as a normalised two-option choice so P(x) and P(not x) agree. How to word it, threshold it and label it.
-
LLM image classification in Python with probabilities
Classify images with a vision LLM from Python: paths, bytes or URLs in images=, typed answers with probabilities, and the privacy and size limits.
-
An LLM decision tree in one request with @key branches
Branch on an earlier LLM answer inside one request: @key in when reads an answer, follow-ups run only on the branch taken, and skips cascade to dependents.
-
LLM classification with many classes: up to 255 options
How an LLM Choice question behaves as labels grow: descriptions, the escape option, and the switch to verbal mode with the top 5 labels past 20 options.
-
Calibrate only when it helps: a held-out gate for LLMs
Calibration can make LLM probabilities worse. How a five-fold held-out gate on log-loss and Brier decides when to apply it, with a real CommonsenseQA case.
-
LLM calibration explained: when 0.9 means 90%
What LLM calibration means, how ECE and Brier measure it, and how Curva fits a calibrator from 30 feedback labels, with held-out before and after numbers.
-
An LLM audit log that never stores the input
What Curva's audit log records for every LLM decision, what it never keeps (the state, the images, provider errors), and how to page through it.
-
Install Curva with pip, npm or Docker and check it runs
Three ways to install Curva, what each gives you, which to pick, and three checks that prove it runs: curva --version, GET /health and a first decision.
-
How Curva works: from state to calibrated answer
The path of one Curva request: rules, cache, fenced prompt, two option orders, logprobs or verbal reading, calibration, coverage set, abstain and audit.
-
Few-shot examples for LLM classification, up to 10
Add up to 10 labeled examples per question, written like feedback labels and remapped when debiasing reverses the options. Which ones to pick.
-
The fair coin test: can an LLM say 50%?
A coin flip has one honest probability. One published probe got 0.92 from a typed-decision model; Curva gave 0.494. What the test shows and its limits.