Python LLM API error handling with CurvaError
Map each CurvaError subclass to the fix: 422 names the bad question, 429 carries retry_after, 502 splits model_error from model_unavailable.
Python LLM API error handling with Curva comes down to one base class and five subclasses. Every failure the client doesn't retry raises a CurvaError with status, type and message, matching the HTTP error the server sent. Catch InvalidRequestError to fix a bad question (a 422 names it), RateLimitError to wait (it carries retry_after), and ModelError for provider failures, where model_error means retry later and model_unavailable means fix the model id. A status of 0 means the server was never reached. And 429 and 5xx responses are retried for you first, so most transient failures never reach your except block.
Python LLM API error handling starts with one base class: CurvaError
from curva import CurvaError
try:
client.decide(state, questions)
except CurvaError as e:
print(e.status, e.type, e.message)The three attributes do different jobs. status is the HTTP status, or 0 when there was no response. type is a fixed word from the server, such as invalid_request, rate_limited or model_unavailable, and it is what your code should branch on. message is for people: it names the question or image at fault, and it is safe to log because provider error details stay in the server log and are never sent to callers.
Log d.request_id on success and keep it on failure where you can: it is the server's x-request-id, the same id that appears in the server's log line for the request.
The five subclasses: AuthError, InvalidRequestError, NotFoundError, RateLimitError, ModelError
| Class | Status | `type` | What to do |
|---|---|---|---|
AuthError | 401 | unauthorized | Send a valid key: CURVA_API_KEY or api_key= |
InvalidRequestError | 400, 422 | invalid_json, invalid_request | Fix the request; the message names the question |
NotFoundError | 404 | not_found | Check the decision id and question key |
RateLimitError | 429 | rate_limited | Wait retry_after seconds, then retry |
ModelError | 502, 504 | model_error, model_unavailable, timeout | Retry later, or fix the model id |
Statuses without a subclass, such as 403 for a key bound to another project or 413 for a body over the server's limit, raise the base CurvaError. Catch subclasses first and the base class last:
from curva import CurvaError, RateLimitError, InvalidRequestError, ModelError
try:
d = client.decide(state, QUESTIONS, project="support")
except InvalidRequestError as e:
log.error("bad question: %s", e.message) # don't retry: fix the question
raise
except RateLimitError as e:
requeue(item, delay=e.retry_after)
except ModelError as e:
if e.type == "model_unavailable":
alert("model id or access is wrong: %s" % e.message)
requeue(item, delay=60)
except CurvaError as e:
log.error("curva %s %s: %s", e.status, e.type, e.message)
raiseStatus 0: curva.local() not started or a wrong CURVA_BASE_URL
A CurvaError with status 0 means the server was unreachable. No HTTP response came back at all. The usual causes:
Curva()is pointed at the wrong place. It readsCURVA_BASE_URLand defaults tohttp://localhost:7777.- No server is running.
Curva()never starts one;curva.local()and the module-levelcurva.decidedo. - The private server didn't come up.
curva.local()waitsstartup_timeoutseconds, 15 by default, for its server to start.
One related error is raised before any request: with no provider key at all and no CURVA_MODEL, the module-level curva.decide raises an error naming the variables to set.
Dropped connections are retried along with 429 and 5xx, so a status 0 that reaches you has already failed max_retries times.
422 invalid_request names the question: option counts, levels, mode conflicts
A 422 means the request was valid JSON but something in it is not allowed. The message names the question or image, so log it in full. Common causes in Python code:
- too few or too many options on a Choice, or too few levels on a Score;
- more than 64 questions in one request, or a state over 150,000 characters;
- bad
max_length,minormaxon an extraction question; mode="logprobs"with a Text, Number or Integer question, or with a Choice over 20 options;- an unknown
configname, an unknown operator inwhenorrules, or adepends_oncycle; privacy="strict"on a named provider whose URL is not on this machine.
A 400 invalid_json is the other InvalidRequestError: the body was not JSON, or named an unknown mode or question type. Neither should be retried. The same request will fail the same way.
502 model_error vs model_unavailable: retry later or fix the model id
Both arrive as ModelError with status 502, and type tells them apart.
- **
model_error**: every model in the chain failed after retries. The provider was down or overloaded. Retry later, or givemodela list so a fallback chain moves to the next model on a 429 or server error. - **
model_unavailable**: the model doesn't exist, was retired, or isn't open to your provider account. Providers rename models often. Retrying won't help; check the model id and your access.
A 504 timeout is also a ModelError: the request took longer than the 120 s deadline.
404 on feedback: unknown decision or question key
client.feedback(decision_id, key, label) raises NotFoundError for an unknown decision or an unknown question key. Two cases surprise people because the decision and key both exist:
- the question was answered by a rule, which is not model output and is not stored for calibration;
- the question was skipped by its
when, so there is no answer to correct.
Check the answer's rule field and the skipped flag before sending feedback. A label that isn't a valid answer, such as a key not among the options, gets 422 and arrives as InvalidRequestError.
What never reaches your except block because the client retried it
The client retries before it raises:
from curva import Curva
client = Curva(timeout=30.0, max_retries=3) # retries 429, 5xx and dropped connectionsBy default, statuses 429, 500, 502, 503 and 504 and dropped connections are retried up to max_retries times, honouring Retry-After. Only what is still failing after that raises. So a RateLimitError in your logs means a limit that lasted through every retry: a key's requests-per-minute limit, the provider's rate limit, or the server's daily budget, which resets at 00:00 UTC. Lower your concurrency or raise the limit rather than adding another retry loop on top.
For batches, AsyncCurva.decide_many(..., return_exceptions=True) puts each failure in its slot instead of raising the first one. Check each result with isinstance(result, CurvaError) and requeue only the failed rows.
Next steps
The docs cover the classes in the Python SDK reference and the statuses in the HTTP API reference. For every status across all clients, read the Curva troubleshooting guide. Tune retries in LLM client retries and timeouts in Python, and handle batches in async LLM classification in Python. Start from the Python LLM classification tutorial.