Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 2 of 18-
Extract numbers from text with an LLM and check the bounds
Number and Integer questions return a typed value checked against min and max, with a confidence. What a bad reply becomes, and why sums stay in code.
-
Expected calibration error (ECE) for LLM classifiers
How expected calibration error is computed, worked by hand on 100 answers, and why an LLM can match on accuracy yet lose on ECE, with measured numbers.
-
Detect display name spoofing in emails
The sender_mismatch check: the display name claims one organisation but the address or reply-to is another domain. Why it is a separate question.
-
Score customer frustration on a scale you define
Use an ordered Score instead of positive or negative: define the levels, read the expected level (0.65 sits between calm and annoyed), and act on it.
-
Curva troubleshooting: every error code and its fix
Every Curva error in one place: status, type, the SDK exception it raises, the usual causes and the fix, plus surprises that are not errors at all.
-
curva.local() vs Curva(): running the server from Python
Three ways to get a Curva client in Python: module-level decide, curva.local() with its own server, or Curva() for a shared one. What each does.
-
Content moderation API with calibrated confidence
The content-moderation recipe: a violation category, a severity score, checks for minors and targeted harassment, and a review queue at 0.85.
-
Conformal prediction sets for LLM classification
Set coverage=0.95 and each answer carries a set of labels holding the right one at least 95% of the time. How split conformal works and what it needs.
-
When conformal guarantees fail: exchangeability for LLMs
Conformal prediction sets need labels that look like your traffic. How review-only labels and shifting inputs break coverage, and how to catch it early.
-
Conditional LLM questions: skip what does not apply
Add when to a question and it is asked only when the state matches. Skipped questions are never sent, stored or paid for, and calibration survives.
-
Which LLMs return logprobs? Check with curva spike
curva spike tests whether a model returns usable label probabilities from logprobs. Run it on free models, a named provider or your own served model.
-
Map customer messages to chargeback reason codes
A Choice over your chargeback reason codes with descriptions per code, the escape option on, and abstain so disputed cases reach a person.
-
Bias scaling: fix an LLM that always picks one answer
When an LLM favours one class, temperature scaling cannot reorder answers. Bias scaling adds a per-answer offset; measured on AITA and Yelp, held out.
-
Batch LLM classification of a JSONL file with curva map
Classify thousands of records from the command line with any LLM: typed answers per line, rate limits respected, and a rerun resumes where it stopped.
-
Classify banking intents with 77 options
Build a 77-intent banking classifier with one Choice: why over 20 options Curva switches to verbal mode, and what Gemini scored on BANKING77 (n = 120).
-
Triage abuse reports with a review branch
Route abuse reports by type and severity with min_confidence, so unsure reports go straight to a person and their verdicts calibrate the model.
-
Abstain or a coverage set? Selective classification for LLMs
Abstain checks the top answer; a conformal set checks every option. A decision table by question type and the routing rule that combines both.
-
Reduce LLM classification cost: every lever, measured
Every way to cut the cost of LLM classification in one place: rules, when, the cache, debias auto, cascades, prompt caching, local models and spend caps.
-
Python LLM classification with confidence and feedback
A start-to-finish Python tutorial: typed LLM labels with a probability each, abstain, feedback, calibration reports, async batches and errors.
-
Phishing detection with an LLM and honest probabilities
Phishing detection with an LLM that returns P(phishing), tactics, a risk level and a sender check, measured on PhishNChips with its limits.
-
n8n AI routing with a Needs review branch
Route n8n items with an LLM: one branch per answer, a Needs review branch for unsure ones, and a Feedback node that calibrates on your data.
-
LLM position bias and how order debiasing cancels it
LLMs favour options by their position in a list. How asking in two orders and averaging cancels it, what it costs, and when debias auto skips a call.
-
LLM decisions over HTTP: one call, any tool
Get typed LLM decisions from any language or no-code tool with two HTTP calls: decide and feedback. Request, response, routing order, errors.
-
LLM classification confidence scores you can act on
How to get confidence scores from LLM classification that mean something: a probability per label, debiasing, calibration on your labels and abstain.