How to word LLM classification questions: recipe lessons

Wording lessons from Curva's shipped recipes: judge by need, define the no case, carve out false positives, and freeze wording before collecting labels.

Good LLM classification instructions wording comes down to a few habits, and Curva's shipped recipes show each one in real text. Tell the model to judge by what the person needs, not the words they use. Say what counts as no. Name the false positives you want excluded. Make "unknown" a real option instead of forcing a guess. Describe every option, including the dull one. Then test the wording on logged traffic and freeze it before you collect feedback, because rewording a question starts its calibration over. This post reads the recipe files as worked examples of each habit.

The recipes are built into the curva binary. curva recipe list names them, and curva recipe show support-triage > questions.json prints one as a questions file you can edit.

Judge by what is needed, not the words used: the support-triage team question

The support-triage recipe's team question reads:

Which team should handle this support ticket? Judge by what the customer needs done, not by the words they happen to use.

The second sentence is the point. Tickets are full of misleading keywords. "I can't log in to see my invoice" mentions an invoice, and a keyword-minded model sends it to billing. What the customer needs done is to get into their account, so it belongs with account. The instruction tells the model which reading to prefer when keywords and intent disagree.

Use the same move wherever your categories have tempting surface cues: classify by the action required or the outcome wanted, not by vocabulary.

Say what counts as no: refund_requested and phishing

Yes/no questions fail at the edges, and the edge is usually a near miss on the no side. The recipes name it:

json
  "refund_requested": {
    "type": "noul",
    "instructions": "The customer explicitly asks for money back (a refund, chargeback or credit). Complaining about a charge without asking for money back is no."
  },

Two details do the work. "Explicitly" sets the bar for yes, and the parenthesis lists what counts as money back. The last sentence names the most common false positive, an angry message about a charge, and rules it out.

The phishing recipe does the same: its phishing question describes the attack (getting the reader to reveal credentials, pay money, open an attachment or click a link under false pretences) and ends with "Legitimate marketing and real notifications are no." Without that sentence, every newsletter with a "click here" button looks suspicious.

Carve out the false positives: quoting is not violating

The content-moderation recipe's violation question carries two rules in its instructions:

Does this user-generated content break a content policy? Pick the most serious category that applies. Quoting, reporting on or condemning harmful content is not itself a violation.

"Pick the most serious category that applies" settles ties: a post with spam and a threat should come back as violence, not spam. "Quoting, reporting on or condemning harmful content is not itself a violation" carves out the biggest false-positive class in moderation. A news post about a threat contains the words of a threat. Without the carve-out, a model reading for keywords flags it.

When you write your own, look at the items your current system gets wrong in a confident way. Most of them share a pattern you can name in one sentence.

Make unknown a real option, then set escape false

The lead-qualification recipe's timeline question:

json
  "timeline": {
    "type": "choice",
    "instructions": "When does the lead expect to buy or start? Use only what they state or clearly imply.",
    "options": {
      "now": "within a month",
      "quarter": "within about three months",
      "later": "more than three months away",
      "unknown": "not stated"
    },
    "escape": false
  },

"Use only what they state or clearly imply" stops the model from inferring a timeline from tone. And unknown with the description "not stated" makes the honest answer a first-class option. Because the options now cover every case, the recipe turns off Curva's default escape option, none_of_these, with "escape": false. Every lead gets one of the four answers, and "not stated" can be counted, routed and reported like the others.

The rule of thumb: if "we can't tell" is a normal outcome for your data, give it its own option. If it is an exception, leave the escape option on and route none_of_these to a person.

Descriptions on every option, including the boring one

Option descriptions are where most of the policy lives. A few patterns from the recipes:

Recipe optionDescription
moderation none"allowed: ordinary content, including criticism, strong opinions and mild profanity"
triage account"login, password reset, two-factor, profile, account access or deletion"
triage urgency level High"the customer cannot use the product, or money is being lost right now"
phishing suspicious_linka link whose text and destination differ, a look-alike domain, or a login page on a domain unrelated to the sender

The first is the important one. The "nothing to see here" option is where false positives go to be caught, so it needs the clearest description of all: criticism and mild profanity are allowed, and saying so keeps them out of harassment. Score levels work the same way. "High" on its own means different things to different readers; "cannot use the product, or money is being lost right now" means one thing.

Lists of concrete examples inside a description help the model more than abstract definitions do. The suspicious_link description even names hosting-panel ports such as :2083, :2087 and :2096, because a login page on such a port and an unrelated domain is a pattern worth spelling out.

Test the classification instructions wording with curva shadow, then freeze it

Recipes are a starting point: edit the wording and options for your domain. Then test before anything depends on it. curva shadow replays logged traffic through your questions and compares the answers with what your current system decided:

bash
curva shadow traffic.jsonl -q questions.json -o shadow.jsonl

The report lists the most confident disagreements first. Read them: each one is either a bug in your old system or a gap in your new wording. Fix the wording, rerun, and repeat until the disagreements are ones you accept. Few-shot examples, up to 10 per question, are the next tool when a boundary stays fuzzy.

Then freeze it. Calibration belongs to a fingerprint of the question's type, wording and options. Rewording a question starts a fresh calibration, with no labels and no coverage guarantee, so settle the wording before you collect feedback. Small edits count too: a reworded question is a new question.

Next steps

The docs list every pack in recipes. For the packs themselves, read Curva recipes. For the yes/no habit in depth, see LLM yes or no questions with a real probability, and for the product overview, what is Curva.