Classify banking intents with 77 options
Build a 77-intent banking classifier with one Choice: why over 20 options Curva switches to verbal mode, and what Gemini scored on BANKING77 (n = 120).
Banking intent classification with an LLM can use one Choice question with all 77 intents as options. Curva accepts up to 255 options, adds a none_of_these escape option, and returns one intent with a probability for each. Over 20 options it reads those probabilities in verbal mode, because providers return at most 20 logprobs. On the public BANKING77 set, Gemini flash-lite scored 79.8% (95% range 73% to 87%, n = 120, 2026-10-01), which matches Jev's published 75.3% within the margin. This post covers the mechanics, the descriptions, the measured numbers and the cost.
In plain words. A big intent list fits in one Choice. Past 20 options the probabilities come from the model's JSON, not token logprobs, and long option lists cost more per call.
Banking intent classification: 77 options in one Choice
A Choice picks exactly one option. Options map a key to a description, and the description may be empty. The order is kept. A Choice takes 2 to 255 options, so a full banking intent list fits in one question:
from curva import Curva, Choice
client = Curva()
INTENTS = {
"card_arrival": "asks when a new card will arrive",
"card_not_working": "a card is declined or not accepted",
"lost_or_stolen_card": "reports a lost or stolen card",
"pending_top_up": "a top-up shows as pending",
"top_up_failed": "a top-up was rejected or failed",
# ... one entry per intent, 77 in all
}
q = {"intent": Choice("What does the customer want?", INTENTS, min_confidence=0.8)}
d = client.decide({"message": "My card got declined at the shop twice today"}, q, project="banking")
print(d["intent"].choice, d["intent"].confidence)Curva also adds a none_of_these option by default, so a message that fits no intent isn't forced into one. In banking, that catches the complaint about branch opening hours that your list never planned for.
The question is asked as one request. If you also want an urgency Score or a "customer is upset" yes/no, add them to the same request; up to 64 questions are answered together.
Over 20 options: verbal mode
Curva reads probabilities in one of two modes. In logprobs mode the model answers with one label token per question, and Curva reads each label's probability from the token log-probabilities. In verbal mode the model returns a JSON object limited to the declared labels, with a probability for each.
Providers return at most 20 logprobs per position. A 77-option question can't be read that way, so Choices with more than 20 options are always answered in verbal mode. The changelog adds that these are answered with the top 5 labels. Setting mode: logprobs on such a question gets a 422.
Three things follow from that:
- Any model works, including models that don't return logprobs at all.
- The answer is the model's own stated distribution, so calibration matters more. A raw 0.95 on a verbal answer is a claim, not a measurement.
- The calibrator is temperature scaling. Bias scaling, which adds a per-answer offset for a model that favours one answer, is used only for questions with up to 20 options.
If the flat list is too much for your model, split it. A first Choice picks a category (cards, transfers, top-ups, account), and a second question that uses depends_on picks the intent within it, seeing the first answer. Each stage is one model call, at most 8 stages per request.
Writing descriptions for close intents
BANKING77 is hard because many intents are near twins: a declined card against a card that doesn't work at all, a pending top-up against a failed one, a transfer that hasn't arrived against one that was declined. The descriptions are what the model reads, so they carry the distinctions.
- Describe what the customer says, not the bank's internal process. "A top-up shows as pending" beats "top-up in settlement".
- Name the deciding detail in each twin. Pending against failed, declined against not arriving.
- Keep keys readable. The model sees them next to the descriptions.
- Don't paste examples into descriptions. Use the question's few-shot
examplesinstead, up to 10(state, label)pairs.
Models also tend to favour options by position, and a long list gives position bias more room. Order debiasing, on by default, asks the question twice with the options in original and reversed order and averages the two.
Measured on BANKING77
Curva's public benchmark runs use the BANKING77 set with 77 options, debiasing on, and samples balanced across intents. Jev's 75.3% is published by others (sanand0 llmevals, 77 items) on a different sample, so compare with care.
| Model | n | Accuracy, natural mix (95% range) | vs Jev 75.3% | ECE raw | ECE after calibration (held out) | p50 |
|---|---|---|---|---|---|---|
@gemini/gemini-flash-lite-latest | 120 | 79.8% (73%–87%) | matches | 0.097 | 0.097 | 1362 ms |
@groq/qwen/qwen3.8-27b | 69 | 74.7% (64%–85%) | matches | 0.252 | 0.157 | 351 ms |
All rows are from 2026-10-01. "After calibration" means each half of the rows was calibrated by a fit on the other half, which is what you get once feedback flows in. Gemini's raw probabilities were already well calibrated, so calibration left them alone. Groq's were overconfident, and calibration brought ECE from 0.252 to 0.157.
The early paid runs, at n = 20 each and dated 2026-10-01, are too small to rank: gpt-4.1-mini 75.0%, Claude Haiku 4.5 65.0%, gpt-4o-mini 60.0%, and gpt-4.1-nano 50.0%, which is a loss against Jev's number. At n = 20 the range is about plus or minus 20 points. Treat these as a reason to test your own model, not as a verdict.
Gemini was accurate but slow here: a median of 1,362 ms per decision. Groq answered in 351 ms. If you need fast answers on a big intent list, measure a fast provider on your own labels.
Cost: the most expensive set we ran
Long option lists mean long prompts, and with debiasing on each decision is two calls. BANKING77 was the most expensive of the 8 public sets in the paid runs. Measured from cost_usd, 20 decisions per model on the set, 2026-09-30, debiasing on:
| Model | Cost per 1,000 decisions, BANKING77 | All 8 sets |
|---|---|---|
@openai/gpt-4.1-nano | $0.294 | $0.078 |
@openai/gpt-4o-mini | $0.441 | $0.118 |
@openai/gpt-4.1-mini | $1.17 | $0.314 |
@anthropic/claude-haiku-4-5-20251001 | $4.78 | $1.64 |
| Gemini flash-lite, Groq qwen3.8-27b | $0 within free-tier limits | $0 |
Figures are truncated, not rounded up. To bring the cost down, answer the obvious intents with rules (a message containing "stolen" can go straight to the lost card flow), let the decision cache answer exact repeats, and try debias: "auto", which drops the second call once a question has shown no position bias.
Next steps
Read the questions reference in the docs and install with pip install curva-ai. For the full benchmark write-up, see the BANKING77 LLM benchmark. For the general problem of big label sets, read LLM classification with many classes. A related fine-grained list is chargeback reason code classification, and more ideas are in LLM classification use cases.