<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel>
<title>Tarkova Blog</title><link>https://www.tarkova.com/blog/</link><description>Practical reads on semantic caching, agent memory, LLM cost and LLM classification.</description><language>en</language>
<atom:link href="https://www.tarkova.com/rss.xml" rel="self" type="application/rss+xml"/>
<item><title>How to word LLM classification questions: recipe lessons</title><link>https://www.tarkova.com/blog/write-llm-classification-instructions/</link><guid>https://www.tarkova.com/blog/write-llm-classification-instructions/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Wording lessons from Curva&#39;s shipped recipes: judge by need, define the no case, carve out false positives, and freeze wording before collecting labels.</description></item>
<item><title>Switch from Jev to Curva: a step-by-step guide</title><link>https://www.tarkova.com/blog/switch-from-jev/</link><guid>https://www.tarkova.com/blog/switch-from-jev/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva vs jev</category><description>Move a Jev integration to Curva: /v1/systemone, the criteria shape, two environment variables, a shadow test, and the differences to expect.</description></item>
<item><title>Support ticket triage with calibrated confidence</title><link>https://www.tarkova.com/blog/support-ticket-triage-ai/</link><guid>https://www.tarkova.com/blog/support-ticket-triage-ai/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>The support-triage recipe end to end: route confident tickets, queue unsure ones, answer obvious ones with rules, and learn from agents&#39; corrections.</description></item>
<item><title>Stop parsing LLM text for decisions</title><link>https://www.tarkova.com/blog/stop-parsing-llm-text/</link><guid>https://www.tarkova.com/blog/stop-parsing-llm-text/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Decisions need types and probabilities, not prose. Why parsing a model&#39;s sentence fails quietly, and what a typed, calibrated decision gives you instead.</description></item>
<item><title>State vs instructions: framing an LLM decision request</title><link>https://www.tarkova.com/blog/state-vs-instructions-llm/</link><guid>https://www.tarkova.com/blog/state-vs-instructions-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>What belongs in the state and what belongs in the question when you ask an LLM for a decision, and which numbers to compute in code before you ask.</description></item>
<item><title>Send feedback to an LLM classifier from Python</title><link>https://www.tarkova.com/blog/send-feedback-llm-classifier-python/</link><guid>https://www.tarkova.com/blog/send-feedback-llm-classifier-python/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Send true labels to an LLM classifier from Python: store the decision id, use the right label per question type, and avoid 404 and 422 errors.</description></item>
<item><title>Classify support tickets that include screenshots</title><link>https://www.tarkova.com/blog/screenshot-ticket-classification/</link><guid>https://www.tarkova.com/blog/screenshot-ticket-classification/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>Send a ticket&#39;s screenshot to a vision model next to its text: is an error visible, which part of the app. Images inline, up to 8 per request.</description></item>
<item><title>Reliability diagrams: plot LLM confidence vs accuracy</title><link>https://www.tarkova.com/blog/reliability-diagram-llm/</link><guid>https://www.tarkova.com/blog/reliability-diagram-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Read a reliability diagram for an LLM classifier from Curva&#39;s reliability bins or dashboard, and spot overconfident, underconfident and leaning models.</description></item>
<item><title>Detect refund requests, not the word refund</title><link>https://www.tarkova.com/blog/refund-request-detection/</link><guid>https://www.tarkova.com/blog/refund-request-detection/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>A yes/no question that separates &quot;I was charged twice&quot; from &quot;please refund me&quot;: the recipe&#39;s wording, the probability it returns and how to calibrate it.</description></item>
<item><title>Plain vs full LLM answers in Python: the hidden 0.5</title><link>https://www.tarkova.com/blog/python-plain-vs-full-llm-answers/</link><guid>https://www.tarkova.com/blog/python-plain-vs-full-llm-answers/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>d.refund is True whenever P(yes) is at least 0.5. When plain answers are fine, when to read the full answer, and the id and model name clash.</description></item>
<item><title>Python LLM API error handling with CurvaError</title><link>https://www.tarkova.com/blog/python-llm-api-error-handling/</link><guid>https://www.tarkova.com/blog/python-llm-api-error-handling/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Map each CurvaError subclass to the fix: 422 names the bad question, 429 carries retry_after, 502 splits model_error from model_unavailable.</description></item>
<item><title>Product catalog tagging with multi-label LLM questions</title><link>https://www.tarkova.com/blog/product-tagging-ai/</link><guid>https://www.tarkova.com/blog/product-tagging-ai/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>Tag products with a Multi question: an independent probability per tag, a threshold you tune, and up to 20 tags per question across a whole catalog.</description></item>
<item><title>Pin an LLM prompt version so answers don&#39;t shift</title><link>https://www.tarkova.com/blog/pin-llm-prompt-version/</link><guid>https://www.tarkova.com/blog/pin-llm-prompt-version/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>A pinned config freezes the prompt template, mode, debias and model so upgrades don&#39;t move your answers. The pins, curva-latest, and when to recalibrate.</description></item>
<item><title>Ollama classification with probabilities, fully local</title><link>https://www.tarkova.com/blog/ollama-llm-classification/</link><guid>https://www.tarkova.com/blog/ollama-llm-classification/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Run typed LLM classification on your own machine with Ollama: no key, no network hop, privacy strict allowed, and a cost of $0 per call.</description></item>
<item><title>None of the above: an escape option for LLM classifiers</title><link>https://www.tarkova.com/blog/none-of-the-above-llm/</link><guid>https://www.tarkova.com/blog/none-of-the-above-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>A classifier forced to pick a label will, confidently. How a none_of_these option works, why it is on by default, and how to route it to a person.</description></item>
<item><title>Logprobs vs verbal confidence: which LLM number to trust</title><link>https://www.tarkova.com/blog/logprobs-vs-verbal-confidence/</link><guid>https://www.tarkova.com/blog/logprobs-vs-verbal-confidence/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Token logprobs or a stated probability per label? How each mode works, which one auto picks, and the probe results that changed Curva&#39;s default model.</description></item>
<item><title>LLM yes or no questions with a real probability of yes</title><link>https://www.tarkova.com/blog/llm-yes-no-probability/</link><guid>https://www.tarkova.com/blog/llm-yes-no-probability/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>A Noul question returns one number, P(yes), asked as a normalised two-option choice so P(x) and P(not x) agree. How to word it, threshold it and label it.</description></item>
<item><title>LLM image classification in Python with probabilities</title><link>https://www.tarkova.com/blog/llm-image-classification-python/</link><guid>https://www.tarkova.com/blog/llm-image-classification-python/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Classify images with a vision LLM from Python: paths, bytes or URLs in images=, typed answers with probabilities, and the privacy and size limits.</description></item>
<item><title>An LLM decision tree in one request with @key branches</title><link>https://www.tarkova.com/blog/llm-decision-tree-one-request/</link><guid>https://www.tarkova.com/blog/llm-decision-tree-one-request/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Branch on an earlier LLM answer inside one request: @key in when reads an answer, follow-ups run only on the branch taken, and skips cascade to dependents.</description></item>
<item><title>LLM classification with many classes: up to 255 options</title><link>https://www.tarkova.com/blog/llm-classification-many-classes/</link><guid>https://www.tarkova.com/blog/llm-classification-many-classes/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>How an LLM Choice question behaves as labels grow: descriptions, the escape option, and the switch to verbal mode with the top 5 labels past 20 options.</description></item>
<item><title>Calibrate only when it helps: a held-out gate for LLMs</title><link>https://www.tarkova.com/blog/llm-calibration-held-out-gate/</link><guid>https://www.tarkova.com/blog/llm-calibration-held-out-gate/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Calibration can make LLM probabilities worse. How a five-fold held-out gate on log-loss and Brier decides when to apply it, with a real CommonsenseQA case.</description></item>
<item><title>LLM calibration explained: when 0.9 means 90%</title><link>https://www.tarkova.com/blog/llm-calibration-explained/</link><guid>https://www.tarkova.com/blog/llm-calibration-explained/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>What LLM calibration means, how ECE and Brier measure it, and how Curva fits a calibrator from 30 feedback labels, with held-out before and after numbers.</description></item>
<item><title>An LLM audit log that never stores the input</title><link>https://www.tarkova.com/blog/llm-audit-log/</link><guid>https://www.tarkova.com/blog/llm-audit-log/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>What Curva&#39;s audit log records for every LLM decision, what it never keeps (the state, the images, provider errors), and how to page through it.</description></item>
<item><title>Install Curva with pip, npm or Docker and check it runs</title><link>https://www.tarkova.com/blog/install-curva/</link><guid>https://www.tarkova.com/blog/install-curva/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Three ways to install Curva, what each gives you, which to pick, and three checks that prove it runs: curva --version, GET /health and a first decision.</description></item>
<item><title>How Curva works: from state to calibrated answer</title><link>https://www.tarkova.com/blog/how-curva-works/</link><guid>https://www.tarkova.com/blog/how-curva-works/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva engineering</category><description>The path of one Curva request: rules, cache, fenced prompt, two option orders, logprobs or verbal reading, calibration, coverage set, abstain and audit.</description></item>
<item><title>Few-shot examples for LLM classification, up to 10</title><link>https://www.tarkova.com/blog/few-shot-examples-llm-classification/</link><guid>https://www.tarkova.com/blog/few-shot-examples-llm-classification/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Add up to 10 labeled examples per question, written like feedback labels and remapped when debiasing reverses the options. Which ones to pick.</description></item>
<item><title>The fair coin test: can an LLM say 50%?</title><link>https://www.tarkova.com/blog/fair-coin-llm-test/</link><guid>https://www.tarkova.com/blog/fair-coin-llm-test/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>A coin flip has one honest probability. One published probe got 0.92 from a typed-decision model; Curva gave 0.494. What the test shows and its limits.</description></item>
<item><title>Extract numbers from text with an LLM and check the bounds</title><link>https://www.tarkova.com/blog/extract-numbers-from-text-llm/</link><guid>https://www.tarkova.com/blog/extract-numbers-from-text-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Number and Integer questions return a typed value checked against min and max, with a confidence. What a bad reply becomes, and why sums stay in code.</description></item>
<item><title>Expected calibration error (ECE) for LLM classifiers</title><link>https://www.tarkova.com/blog/expected-calibration-error-llm/</link><guid>https://www.tarkova.com/blog/expected-calibration-error-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>How expected calibration error is computed, worked by hand on 100 answers, and why an LLM can match on accuracy yet lose on ECE, with measured numbers.</description></item>
<item><title>Detect display name spoofing in emails</title><link>https://www.tarkova.com/blog/display-name-spoofing-detection/</link><guid>https://www.tarkova.com/blog/display-name-spoofing-detection/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>The sender_mismatch check: the display name claims one organisation but the address or reply-to is another domain. Why it is a separate question.</description></item>
<item><title>Score customer frustration on a scale you define</title><link>https://www.tarkova.com/blog/customer-frustration-score/</link><guid>https://www.tarkova.com/blog/customer-frustration-score/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>Use an ordered Score instead of positive or negative: define the levels, read the expected level (0.65 sits between calm and annoyed), and act on it.</description></item>
<item><title>Curva troubleshooting: every error code and its fix</title><link>https://www.tarkova.com/blog/curva-troubleshooting-guide/</link><guid>https://www.tarkova.com/blog/curva-troubleshooting-guide/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Every Curva error in one place: status, type, the SDK exception it raises, the usual causes and the fix, plus surprises that are not errors at all.</description></item>
<item><title>curva.local() vs Curva(): running the server from Python</title><link>https://www.tarkova.com/blog/curva-local-python-server/</link><guid>https://www.tarkova.com/blog/curva-local-python-server/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Three ways to get a Curva client in Python: module-level decide, curva.local() with its own server, or Curva() for a shared one. What each does.</description></item>
<item><title>Content moderation API with calibrated confidence</title><link>https://www.tarkova.com/blog/content-moderation-api/</link><guid>https://www.tarkova.com/blog/content-moderation-api/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>The content-moderation recipe: a violation category, a severity score, checks for minors and targeted harassment, and a review queue at 0.85.</description></item>
<item><title>Conformal prediction sets for LLM classification</title><link>https://www.tarkova.com/blog/conformal-prediction-llm/</link><guid>https://www.tarkova.com/blog/conformal-prediction-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Set coverage=0.95 and each answer carries a set of labels holding the right one at least 95% of the time. How split conformal works and what it needs.</description></item>
<item><title>When conformal guarantees fail: exchangeability for LLMs</title><link>https://www.tarkova.com/blog/conformal-prediction-assumptions/</link><guid>https://www.tarkova.com/blog/conformal-prediction-assumptions/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Conformal prediction sets need labels that look like your traffic. How review-only labels and shifting inputs break coverage, and how to catch it early.</description></item>
<item><title>Conditional LLM questions: skip what does not apply</title><link>https://www.tarkova.com/blog/conditional-llm-questions-when/</link><guid>https://www.tarkova.com/blog/conditional-llm-questions-when/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Add when to a question and it is asked only when the state matches. Skipped questions are never sent, stored or paid for, and calibration survives.</description></item>
<item><title>Which LLMs return logprobs? Check with curva spike</title><link>https://www.tarkova.com/blog/check-llm-logprobs-support/</link><guid>https://www.tarkova.com/blog/check-llm-logprobs-support/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>curva spike tests whether a model returns usable label probabilities from logprobs. Run it on free models, a named provider or your own served model.</description></item>
<item><title>Map customer messages to chargeback reason codes</title><link>https://www.tarkova.com/blog/chargeback-reason-classification/</link><guid>https://www.tarkova.com/blog/chargeback-reason-classification/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>A Choice over your chargeback reason codes with descriptions per code, the escape option on, and abstain so disputed cases reach a person.</description></item>
<item><title>Bias scaling: fix an LLM that always picks one answer</title><link>https://www.tarkova.com/blog/bias-scaling-llm-calibration/</link><guid>https://www.tarkova.com/blog/bias-scaling-llm-calibration/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>When an LLM favours one class, temperature scaling cannot reorder answers. Bias scaling adds a per-answer offset; measured on AITA and Yelp, held out.</description></item>
<item><title>Batch LLM classification of a JSONL file with curva map</title><link>https://www.tarkova.com/blog/batch-llm-classification-jsonl/</link><guid>https://www.tarkova.com/blog/batch-llm-classification-jsonl/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Classify thousands of records from the command line with any LLM: typed answers per line, rate limits respected, and a rerun resumes where it stopped.</description></item>
<item><title>Classify banking intents with 77 options</title><link>https://www.tarkova.com/blog/banking-intent-classification/</link><guid>https://www.tarkova.com/blog/banking-intent-classification/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>Build a 77-intent banking classifier with one Choice: why over 20 options Curva switches to verbal mode, and what Gemini scored on BANKING77 (n = 120).</description></item>
<item><title>Triage abuse reports with a review branch</title><link>https://www.tarkova.com/blog/abuse-report-triage/</link><guid>https://www.tarkova.com/blog/abuse-report-triage/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>Route abuse reports by type and severity with min_confidence, so unsure reports go straight to a person and their verdicts calibrate the model.</description></item>
<item><title>Abstain or a coverage set? Selective classification for LLMs</title><link>https://www.tarkova.com/blog/abstain-vs-conformal-prediction/</link><guid>https://www.tarkova.com/blog/abstain-vs-conformal-prediction/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>Abstain checks the top answer; a conformal set checks every option. A decision table by question type and the routing rule that combines both.</description></item>
<item><title>Reduce LLM classification cost: every lever, measured</title><link>https://www.tarkova.com/blog/reduce-llm-classification-cost/</link><guid>https://www.tarkova.com/blog/reduce-llm-classification-cost/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Every way to cut the cost of LLM classification in one place: rules, when, the cache, debias auto, cascades, prompt caching, local models and spend caps.</description></item>
<item><title>Python LLM classification with confidence and feedback</title><link>https://www.tarkova.com/blog/python-llm-classification/</link><guid>https://www.tarkova.com/blog/python-llm-classification/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>A start-to-finish Python tutorial: typed LLM labels with a probability each, abstain, feedback, calibration reports, async batches and errors.</description></item>
<item><title>Phishing detection with an LLM and honest probabilities</title><link>https://www.tarkova.com/blog/phishing-detection-llm/</link><guid>https://www.tarkova.com/blog/phishing-detection-llm/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva use cases</category><description>Phishing detection with an LLM that returns P(phishing), tactics, a risk level and a sender check, measured on PhishNChips with its limits.</description></item>
<item><title>n8n AI routing with a Needs review branch</title><link>https://www.tarkova.com/blog/n8n-ai-routing/</link><guid>https://www.tarkova.com/blog/n8n-ai-routing/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Route n8n items with an LLM: one branch per answer, a Needs review branch for unsure ones, and a Feedback node that calibrates on your data.</description></item>
<item><title>LLM position bias and how order debiasing cancels it</title><link>https://www.tarkova.com/blog/llm-position-bias/</link><guid>https://www.tarkova.com/blog/llm-position-bias/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva concepts</category><description>LLMs favour options by their position in a list. How asking in two orders and averaging cancels it, what it costs, and when debias auto skips a call.</description></item>
<item><title>LLM decisions over HTTP: one call, any tool</title><link>https://www.tarkova.com/blog/llm-decisions-over-http/</link><guid>https://www.tarkova.com/blog/llm-decisions-over-http/</guid><pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate><category>curva guides</category><description>Get typed LLM decisions from any language or no-code tool with two HTTP calls: decide and feedback. Request, response, routing order, errors.</description></item>
</channel></rss>