Tutorial

Calling Jev API from Python: a robust client with timeouts, retries and error handling

Build a production-ready Python client for Jev API with requests: connect and read timeouts, exponential backoff for 429 and 5xx, and typed errors for 401 and 402.

A one-line requests.post is enough for a demo. Production code needs more. It should give up on a hung connection instead of waiting forever, retry the failures that are actually temporary, stop immediately on the ones that are not, and refuse to treat a malformed response as a decision. This tutorial builds a small Python client for Jev API that does all of that, using only the requests library.

Everything here is based on how POST https://jev-api.com/v1/systemone behaves: Bearer authentication with a jev_live_ or jev_test_ key, a JSON body of model, state and questions, and a JSON error envelope with a request ID. If you have not made a first call yet, start with the getting started guide.

Know which errors are worth retrying

Every error response has the same shape, and the same request ID is returned in the x-request-id header:

{ "error": { "code": "insufficient_balance", "message": "Balance 0.000000 USD is below the estimated 0.000057 USD charge.", "request_id": "req_…" } }

Status codes fall into two groups. Retrying the first group never helps. Retrying the second group often does:

  • Do not retry. 400 invalid_request means the body is wrong, for example a missing state, an empty questions object or an unknown model. 401 unauthorized means the key is missing, mistyped or revoked (“Invalid or revoked API key.”). 402 insufficient_balance means your prepaid balance is below the estimated cost of the call. Repeating any of these sends the same request into the same failure.
  • Retry with backoff. 429 service_error means the model service is rate limiting. 502 service_error means it could not be reached or returned something unusable. 503 service_error means it is temporarily unavailable. 500 internal_error is an unexpected server error. Connection failures on your side belong here too.

Failed calls are not charged, because credit is deducted only after a call succeeds. A retry after an error status therefore does not double-bill you. Timeouts are different, as the next section explains.

Timeouts, and why read timeouts are special

requests has no default timeout. Without one, a stalled connection can block a worker indefinitely. Pass a (connect, read) tuple. A short connect timeout catches network problems quickly. A longer read timeout leaves room for the model to evaluate your questions.

The two timeouts mean different things. A connect timeout means the request never reached the server, so retrying is safe. A read timeout means the request was sent, but the response did not arrive in time. The server may still have finished the call, recorded it and deducted its cost, and your client has no way to tell. Retrying could pay for the same decision twice. The client below therefore does not retry read timeouts by default. Turn that on only if an occasional duplicate charge is acceptable for your workload. You can check what actually happened in the usage table on your dashboard.

Choosing the values

There is no single right number. A connect timeout of a few seconds is plenty on a healthy network. For the read timeout, measure. Log how long successful calls take with your real payloads, then set the timeout comfortably above the slowest normal call. Calls with large states or many questions take longer than a one-line message with a single noul. If a user is waiting on the decision, a shorter timeout and a fallback, such as sending the item to manual review, is usually better than a long wait.

Why not just mount urllib3's Retry?

requests can retry through an HTTPAdapter configured with urllib3.util.Retry, and for GET requests that is often all you need. It is a poor fit here for two reasons. First, by default urllib3 does not retry POST requests at all, and POST /v1/systemone is the call you care about. Second, even with POST enabled and status_forcelist set, the adapter only knows status codes. It cannot turn a 402 into a “stop the batch” signal, attach the API's request_id to the error, or treat read timeouts differently from connection failures. A small explicit loop keeps those decisions visible in your own code.

The client

Install the one dependency with pip install requests, then save this as jev_client.py:

import os
import random
import time
from typing import Any, Optional

import requests

API_URL = "https://jev-api.com/v1/systemone"
RETRYABLE_STATUSES = {429, 500, 502, 503}


class JevError(Exception):
    """Base error. status is None when no HTTP response was received."""

    def __init__(self, message: str, status: Optional[int] = None,
                 code: Optional[str] = None, request_id: Optional[str] = None):
        super().__init__(message)
        self.status = status
        self.code = code
        self.request_id = request_id


class JevRequestError(JevError):
    """400 invalid_request: fix the payload."""


class JevAuthError(JevError):
    """401 unauthorized: missing, invalid or revoked key."""


class JevBalanceError(JevError):
    """402 insufficient_balance: add credit before retrying."""


class JevServiceError(JevError):
    """429/5xx after retries, network failure, or an unusable response."""


ERROR_CLASSES = {400: JevRequestError, 401: JevAuthError, 402: JevBalanceError}


class JevClient:
    def __init__(self, api_key: Optional[str] = None, *, url: str = API_URL,
                 connect_timeout: float = 5.0, read_timeout: float = 30.0,
                 max_retries: int = 3, backoff_base: float = 0.5,
                 retry_read_timeouts: bool = False):
        key = api_key or os.environ.get("JEV_API_KEY", "")
        if not key.startswith(("jev_live_", "jev_test_")):
            raise ValueError("Set JEV_API_KEY to a jev_live_ or jev_test_ key.")
        self.url = url
        self.timeout = (connect_timeout, read_timeout)
        self.max_retries = max_retries
        self.backoff_base = backoff_base
        self.retry_read_timeouts = retry_read_timeouts
        self.session = requests.Session()
        self.session.headers.update({
            "Authorization": f"Bearer {key}",
            "Content-Type": "application/json",
        })

    def decide(self, state: Any, questions: dict, model: str = "jev-latest") -> dict:
        payload = {"model": model, "state": state, "questions": questions}
        attempt = 0
        while True:
            retry_after = None
            try:
                resp = self.session.post(self.url, json=payload, timeout=self.timeout)
            except requests.exceptions.ReadTimeout as exc:
                # The request was sent; it may still complete and be billed.
                if not self.retry_read_timeouts or attempt >= self.max_retries:
                    raise JevServiceError(f"Read timeout: {exc}") from exc
            except requests.exceptions.ConnectionError as exc:
                if attempt >= self.max_retries:
                    raise JevServiceError(f"Network error: {exc}") from exc
            else:
                if resp.status_code == 200:
                    return self._parse_success(resp, questions)
                error = self._parse_error(resp)
                if resp.status_code not in RETRYABLE_STATUSES or attempt >= self.max_retries:
                    raise error
                retry_after = resp.headers.get("Retry-After")
            time.sleep(self._delay(attempt, retry_after))
            attempt += 1

    def _delay(self, attempt: int, retry_after: Optional[str]) -> float:
        if retry_after:
            try:
                return min(float(retry_after), 60.0)
            except ValueError:
                pass
        # Exponential backoff with jitter: ~0.5s, 1s, 2s ... plus up to backoff_base.
        return self.backoff_base * (2 ** attempt) + random.uniform(0, self.backoff_base)

    @staticmethod
    def _parse_error(resp: requests.Response) -> JevError:
        code = message = request_id = None
        try:
            err = resp.json().get("error") or {}
            code, message, request_id = err.get("code"), err.get("message"), err.get("request_id")
        except (ValueError, AttributeError):
            pass
        request_id = request_id or resp.headers.get("x-request-id")
        cls = ERROR_CLASSES.get(resp.status_code, JevServiceError)
        return cls(message or f"HTTP {resp.status_code}", resp.status_code, code, request_id)

    @staticmethod
    def _parse_success(resp: requests.Response, questions: dict) -> dict:
        request_id = resp.headers.get("x-request-id")
        try:
            data = resp.json()
        except ValueError as exc:
            raise JevServiceError("Response was not JSON.", 200, None, request_id) from exc
        answers = data.get("answers") if isinstance(data, dict) else None
        if not isinstance(answers, dict):
            raise JevServiceError("Response has no answers object.", 200, None, request_id)
        for key, question in questions.items():
            answer = answers.get(key)
            if not isinstance(answer, dict) or answer.get("type") != question.get("type"):
                raise JevServiceError(f"Missing or mismatched answer for {key!r}.", 200, None, request_id)
        data["request_id"] = request_id
        return data

What each part does

  • Key handling. The key comes from the JEV_API_KEY environment variable and is checked for a valid prefix at start-up. The client never prints it. A requests.Session reuses the TCP connection across calls and sets the headers once.
  • Error classes. Callers can catch JevBalanceError or JevAuthError specifically. Both mean a person needs to act, and no amount of retrying will fix them. JevRequestError points to a bug in how you build payloads.
  • Backoff. Each retry waits roughly twice as long as the last one. Random jitter keeps many workers from retrying in lockstep. If a response includes a Retry-After header, the client honours it, capped at 60 seconds.
  • Answer validation. A 200 response is only accepted if every question you asked has an answer of the same type. Your code never has to guess what a missing key means.
  • Request ID. The x-request-id header is copied into the result and into every exception, ready for your logs.

Using the client

Here is the client triaging a support message. The questions are an ordinary dict, so you can keep them in one module and reuse them across calls.

from jev_client import JevClient

QUESTIONS = {
    "department": {
        "type": "choice",
        "instructions": "Which team should handle this message?",
        "criteria": {
            "billing": "Charges, invoices, refunds or subscription changes",
            "technical": "Bugs, outages, errors or integration problems",
            "sales": "Pricing questions or new purchases",
        },
    },
    "is_urgent": {
        "type": "noul",
        "instructions": "Does this message need a reply today?",
    },
}

client = JevClient(read_timeout=20)

result = client.decide(
    state="Our checkout has returned errors since this morning and no orders are coming in.",
    questions=QUESTIONS,
)
department = result["answers"]["department"]
urgent = result["answers"]["is_urgent"]["noul"]
print(department["choice"], round(department["confidence"], 2), round(urgent, 2))
print(result["model"], result["usage"]["input_tokens"], result["request_id"])

Processing a batch without burning credit

Batch jobs are where error handling matters most. When the balance runs out halfway through 10,000 tickets, every remaining call will return 402. The job should stop at the first one, not hammer the API 9,000 more times. The same applies to a revoked key.

import logging

from jev_client import JevAuthError, JevBalanceError, JevClient, JevError

log = logging.getLogger("triage")


def triage_all(tickets, client: JevClient, questions: dict):
    results = {}
    for ticket in tickets:
        try:
            out = client.decide(state=ticket["text"], questions=questions)
        except (JevBalanceError, JevAuthError) as exc:
            log.error("Stopping batch: %s (request_id=%s)", exc, exc.request_id)
            break
        except JevError as exc:
            log.warning("Skipping ticket %s: %s (request_id=%s)", ticket["id"], exc, exc.request_id)
            continue
        results[ticket["id"]] = out["answers"]
        log.info("ticket=%s model=%s input_tokens=%s request_id=%s", ticket["id"],
                 out["model"], out["usage"]["input_tokens"], out["request_id"])
    return results

If you add concurrency, for example with concurrent.futures.ThreadPoolExecutor, keep the number of workers small and grow it slowly. A rising count of 429 responses means you are going faster than the model service accepts. Jev API does not publish a fixed rate limit, so treat 429s as the signal and let the backoff absorb short bursts. Note that requests.Session is not officially thread-safe, so give each worker thread its own JevClient.

Testing the client without spending credit

Retry and error paths are hard to trigger against the live API on demand, so test them with fake responses. unittest.mock can replace session.post. These tests check that a 402 is not retried and that a 503 is:

import json
from unittest import mock

import pytest
import requests

from jev_client import JevBalanceError, JevClient

QUESTION = {"q": {"type": "noul", "instructions": "Is this spam?"}}


def fake_response(status, body, headers=None):
    resp = requests.Response()
    resp.status_code = status
    resp._content = json.dumps(body).encode()
    resp.headers.update(headers or {})
    return resp


def test_402_is_not_retried():
    client = JevClient("jev_test_xxx", max_retries=3)
    body = {"error": {"code": "insufficient_balance", "message": "Balance too low", "request_id": "req_1"}}
    with mock.patch.object(client.session, "post", return_value=fake_response(402, body)) as post:
        with pytest.raises(JevBalanceError) as info:
            client.decide("text", QUESTION)
    assert post.call_count == 1
    assert info.value.request_id == "req_1"


def test_503_is_retried_then_succeeds():
    client = JevClient("jev_test_xxx", max_retries=3, backoff_base=0)
    unavailable = fake_response(503, {"error": {"code": "service_error", "message": "x", "request_id": "req_2"}})
    ok = fake_response(200, {"model": "jev-1.13.0",
                             "answers": {"q": {"type": "noul", "noul": 0.1}},
                             "usage": {"input_tokens": 30, "output_tokens": 5}})
    with mock.patch.object(client.session, "post", side_effect=[unavailable, ok]) as post:
        result = client.decide("text", QUESTION)
    assert post.call_count == 2
    assert result["answers"]["q"]["noul"] == 0.1

Logging and secrets

Log the response model, usage.input_tokens and the request ID for every call. The model tells you which version produced each decision. The token count lets you reconcile your own records with the usage table on your dashboard. Never log the Authorization header or the key itself. If you dump request objects while debugging, strip the header first. If a key does end up somewhere it should not, revoke it on the dashboard and create a new one. Revocation takes effect immediately for new requests.

Next steps

With a client that fails safely, the next improvements are better questions and a cost budget:

  • Docs: account, key and credit setup from start to finish.
  • API reference: the full request and response contract for POST /v1/systemone and GET /v1/models.
  • Pricing: current per-token price and how prepaid credit works.

More tutorials on this blog:

← All posts