Python LLM API error handling with CurvaError

Map each CurvaError subclass to the fix: 422 names the bad question, 429 carries retry_after, 502 splits model_error from model_unavailable.

Python LLM API error handling with Curva comes down to one base class and five subclasses. Every failure the client doesn't retry raises a CurvaError with status, type and message, matching the HTTP error the server sent. Catch InvalidRequestError to fix a bad question (a 422 names it), RateLimitError to wait (it carries retry_after), and ModelError for provider failures, where model_error means retry later and model_unavailable means fix the model id. A status of 0 means the server was never reached. And 429 and 5xx responses are retried for you first, so most transient failures never reach your except block.

Python LLM API error handling starts with one base class: CurvaError

python
from curva import CurvaError

try:
    client.decide(state, questions)
except CurvaError as e:
    print(e.status, e.type, e.message)

The three attributes do different jobs. status is the HTTP status, or 0 when there was no response. type is a fixed word from the server, such as invalid_request, rate_limited or model_unavailable, and it is what your code should branch on. message is for people: it names the question or image at fault, and it is safe to log because provider error details stay in the server log and are never sent to callers.

Log d.request_id on success and keep it on failure where you can: it is the server's x-request-id, the same id that appears in the server's log line for the request.

The five subclasses: AuthError, InvalidRequestError, NotFoundError, RateLimitError, ModelError

ClassStatus`type`What to do
AuthError401unauthorizedSend a valid key: CURVA_API_KEY or api_key=
InvalidRequestError400, 422invalid_json, invalid_requestFix the request; the message names the question
NotFoundError404not_foundCheck the decision id and question key
RateLimitError429rate_limitedWait retry_after seconds, then retry
ModelError502, 504model_error, model_unavailable, timeoutRetry later, or fix the model id

Statuses without a subclass, such as 403 for a key bound to another project or 413 for a body over the server's limit, raise the base CurvaError. Catch subclasses first and the base class last:

python
from curva import CurvaError, RateLimitError, InvalidRequestError, ModelError

try:
    d = client.decide(state, QUESTIONS, project="support")
except InvalidRequestError as e:
    log.error("bad question: %s", e.message)    # don't retry: fix the question
    raise
except RateLimitError as e:
    requeue(item, delay=e.retry_after)
except ModelError as e:
    if e.type == "model_unavailable":
        alert("model id or access is wrong: %s" % e.message)
    requeue(item, delay=60)
except CurvaError as e:
    log.error("curva %s %s: %s", e.status, e.type, e.message)
    raise

Status 0: curva.local() not started or a wrong CURVA_BASE_URL

A CurvaError with status 0 means the server was unreachable. No HTTP response came back at all. The usual causes:

  • Curva() is pointed at the wrong place. It reads CURVA_BASE_URL and defaults to http://localhost:7777.
  • No server is running. Curva() never starts one; curva.local() and the module-level curva.decide do.
  • The private server didn't come up. curva.local() waits startup_timeout seconds, 15 by default, for its server to start.

One related error is raised before any request: with no provider key at all and no CURVA_MODEL, the module-level curva.decide raises an error naming the variables to set.

Dropped connections are retried along with 429 and 5xx, so a status 0 that reaches you has already failed max_retries times.

422 invalid_request names the question: option counts, levels, mode conflicts

A 422 means the request was valid JSON but something in it is not allowed. The message names the question or image, so log it in full. Common causes in Python code:

  • too few or too many options on a Choice, or too few levels on a Score;
  • more than 64 questions in one request, or a state over 150,000 characters;
  • bad max_length, min or max on an extraction question;
  • mode="logprobs" with a Text, Number or Integer question, or with a Choice over 20 options;
  • an unknown config name, an unknown operator in when or rules, or a depends_on cycle;
  • privacy="strict" on a named provider whose URL is not on this machine.

A 400 invalid_json is the other InvalidRequestError: the body was not JSON, or named an unknown mode or question type. Neither should be retried. The same request will fail the same way.

502 model_error vs model_unavailable: retry later or fix the model id

Both arrive as ModelError with status 502, and type tells them apart.

  • **model_error**: every model in the chain failed after retries. The provider was down or overloaded. Retry later, or give model a list so a fallback chain moves to the next model on a 429 or server error.
  • **model_unavailable**: the model doesn't exist, was retired, or isn't open to your provider account. Providers rename models often. Retrying won't help; check the model id and your access.

A 504 timeout is also a ModelError: the request took longer than the 120 s deadline.

404 on feedback: unknown decision or question key

client.feedback(decision_id, key, label) raises NotFoundError for an unknown decision or an unknown question key. Two cases surprise people because the decision and key both exist:

  • the question was answered by a rule, which is not model output and is not stored for calibration;
  • the question was skipped by its when, so there is no answer to correct.

Check the answer's rule field and the skipped flag before sending feedback. A label that isn't a valid answer, such as a key not among the options, gets 422 and arrives as InvalidRequestError.

What never reaches your except block because the client retried it

The client retries before it raises:

python
from curva import Curva

client = Curva(timeout=30.0, max_retries=3)   # retries 429, 5xx and dropped connections

By default, statuses 429, 500, 502, 503 and 504 and dropped connections are retried up to max_retries times, honouring Retry-After. Only what is still failing after that raises. So a RateLimitError in your logs means a limit that lasted through every retry: a key's requests-per-minute limit, the provider's rate limit, or the server's daily budget, which resets at 00:00 UTC. Lower your concurrency or raise the limit rather than adding another retry loop on top.

For batches, AsyncCurva.decide_many(..., return_exceptions=True) puts each failure in its slot instead of raising the first one. Check each result with isinstance(result, CurvaError) and requeue only the failed rows.

Next steps

The docs cover the classes in the Python SDK reference and the statuses in the HTTP API reference. For every status across all clients, read the Curva troubleshooting guide. Tune retries in LLM client retries and timeouts in Python, and handle batches in async LLM classification in Python. Start from the Python LLM classification tutorial.