A one-line requests.post is enough for a demo. Production code needs more. It should give up on a hung connection instead of waiting forever, retry the failures that are actually temporary, stop immediately on the ones that are not, and refuse to treat a malformed response as a decision. This tutorial builds a small Python client for Jev API that does all of that, using only the requests library.
Everything here is based on how POST https://jev-api.com/v1/systemone behaves: Bearer authentication with a jev_live_ or jev_test_ key, a JSON body of model, state and questions, and a JSON error envelope with a request ID. If you have not made a first call yet, start with the getting started guide.
Know which errors are worth retrying
Every error response has the same shape, and the same request ID is returned in the x-request-id header:
{ "error": { "code": "insufficient_balance", "message": "Balance 0.000000 USD is below the estimated 0.000057 USD charge.", "request_id": "req_…" } }
Status codes fall into two groups. Retrying the first group never helps. Retrying the second group often does:
- Do not retry.
400 invalid_requestmeans the body is wrong, for example a missingstate, an emptyquestionsobject or an unknown model.401 unauthorizedmeans the key is missing, mistyped or revoked (“Invalid or revoked API key.”).402 insufficient_balancemeans your prepaid balance is below the estimated cost of the call. Repeating any of these sends the same request into the same failure. - Retry with backoff.
429 service_errormeans the model service is rate limiting.502 service_errormeans it could not be reached or returned something unusable.503 service_errormeans it is temporarily unavailable.500 internal_erroris an unexpected server error. Connection failures on your side belong here too.
Failed calls are not charged, because credit is deducted only after a call succeeds. A retry after an error status therefore does not double-bill you. Timeouts are different, as the next section explains.
Timeouts, and why read timeouts are special
requests has no default timeout. Without one, a stalled connection can block a worker indefinitely. Pass a (connect, read) tuple. A short connect timeout catches network problems quickly. A longer read timeout leaves room for the model to evaluate your questions.
The two timeouts mean different things. A connect timeout means the request never reached the server, so retrying is safe. A read timeout means the request was sent, but the response did not arrive in time. The server may still have finished the call, recorded it and deducted its cost, and your client has no way to tell. Retrying could pay for the same decision twice. The client below therefore does not retry read timeouts by default. Turn that on only if an occasional duplicate charge is acceptable for your workload. You can check what actually happened in the usage table on your dashboard.
Choosing the values
There is no single right number. A connect timeout of a few seconds is plenty on a healthy network. For the read timeout, measure. Log how long successful calls take with your real payloads, then set the timeout comfortably above the slowest normal call. Calls with large states or many questions take longer than a one-line message with a single noul. If a user is waiting on the decision, a shorter timeout and a fallback, such as sending the item to manual review, is usually better than a long wait.
Why not just mount urllib3's Retry?
requests can retry through an HTTPAdapter configured with urllib3.util.Retry, and for GET requests that is often all you need. It is a poor fit here for two reasons. First, by default urllib3 does not retry POST requests at all, and POST /v1/systemone is the call you care about. Second, even with POST enabled and status_forcelist set, the adapter only knows status codes. It cannot turn a 402 into a “stop the batch” signal, attach the API's request_id to the error, or treat read timeouts differently from connection failures. A small explicit loop keeps those decisions visible in your own code.
The client
Install the one dependency with pip install requests, then save this as jev_client.py:
import os
import random
import time
from typing import Any, Optional
import requests
API_URL = "https://jev-api.com/v1/systemone"
RETRYABLE_STATUSES = {429, 500, 502, 503}
class JevError(Exception):
"""Base error. status is None when no HTTP response was received."""
def __init__(self, message: str, status: Optional[int] = None,
code: Optional[str] = None, request_id: Optional[str] = None):
super().__init__(message)
self.status = status
self.code = code
self.request_id = request_id
class JevRequestError(JevError):
"""400 invalid_request: fix the payload."""
class JevAuthError(JevError):
"""401 unauthorized: missing, invalid or revoked key."""
class JevBalanceError(JevError):
"""402 insufficient_balance: add credit before retrying."""
class JevServiceError(JevError):
"""429/5xx after retries, network failure, or an unusable response."""
ERROR_CLASSES = {400: JevRequestError, 401: JevAuthError, 402: JevBalanceError}
class JevClient:
def __init__(self, api_key: Optional[str] = None, *, url: str = API_URL,
connect_timeout: float = 5.0, read_timeout: float = 30.0,
max_retries: int = 3, backoff_base: float = 0.5,
retry_read_timeouts: bool = False):
key = api_key or os.environ.get("JEV_API_KEY", "")
if not key.startswith(("jev_live_", "jev_test_")):
raise ValueError("Set JEV_API_KEY to a jev_live_ or jev_test_ key.")
self.url = url
self.timeout = (connect_timeout, read_timeout)
self.max_retries = max_retries
self.backoff_base = backoff_base
self.retry_read_timeouts = retry_read_timeouts
self.session = requests.Session()
self.session.headers.update({
"Authorization": f"Bearer {key}",
"Content-Type": "application/json",
})
def decide(self, state: Any, questions: dict, model: str = "jev-latest") -> dict:
payload = {"model": model, "state": state, "questions": questions}
attempt = 0
while True:
retry_after = None
try:
resp = self.session.post(self.url, json=payload, timeout=self.timeout)
except requests.exceptions.ReadTimeout as exc:
# The request was sent; it may still complete and be billed.
if not self.retry_read_timeouts or attempt >= self.max_retries:
raise JevServiceError(f"Read timeout: {exc}") from exc
except requests.exceptions.ConnectionError as exc:
if attempt >= self.max_retries:
raise JevServiceError(f"Network error: {exc}") from exc
else:
if resp.status_code == 200:
return self._parse_success(resp, questions)
error = self._parse_error(resp)
if resp.status_code not in RETRYABLE_STATUSES or attempt >= self.max_retries:
raise error
retry_after = resp.headers.get("Retry-After")
time.sleep(self._delay(attempt, retry_after))
attempt += 1
def _delay(self, attempt: int, retry_after: Optional[str]) -> float:
if retry_after:
try:
return min(float(retry_after), 60.0)
except ValueError:
pass
# Exponential backoff with jitter: ~0.5s, 1s, 2s ... plus up to backoff_base.
return self.backoff_base * (2 ** attempt) + random.uniform(0, self.backoff_base)
@staticmethod
def _parse_error(resp: requests.Response) -> JevError:
code = message = request_id = None
try:
err = resp.json().get("error") or {}
code, message, request_id = err.get("code"), err.get("message"), err.get("request_id")
except (ValueError, AttributeError):
pass
request_id = request_id or resp.headers.get("x-request-id")
cls = ERROR_CLASSES.get(resp.status_code, JevServiceError)
return cls(message or f"HTTP {resp.status_code}", resp.status_code, code, request_id)
@staticmethod
def _parse_success(resp: requests.Response, questions: dict) -> dict:
request_id = resp.headers.get("x-request-id")
try:
data = resp.json()
except ValueError as exc:
raise JevServiceError("Response was not JSON.", 200, None, request_id) from exc
answers = data.get("answers") if isinstance(data, dict) else None
if not isinstance(answers, dict):
raise JevServiceError("Response has no answers object.", 200, None, request_id)
for key, question in questions.items():
answer = answers.get(key)
if not isinstance(answer, dict) or answer.get("type") != question.get("type"):
raise JevServiceError(f"Missing or mismatched answer for {key!r}.", 200, None, request_id)
data["request_id"] = request_id
return data
What each part does
- Key handling. The key comes from the
JEV_API_KEYenvironment variable and is checked for a valid prefix at start-up. The client never prints it. Arequests.Sessionreuses the TCP connection across calls and sets the headers once. - Error classes. Callers can catch
JevBalanceErrororJevAuthErrorspecifically. Both mean a person needs to act, and no amount of retrying will fix them.JevRequestErrorpoints to a bug in how you build payloads. - Backoff. Each retry waits roughly twice as long as the last one. Random jitter keeps many workers from retrying in lockstep. If a response includes a
Retry-Afterheader, the client honours it, capped at 60 seconds. - Answer validation. A 200 response is only accepted if every question you asked has an answer of the same type. Your code never has to guess what a missing key means.
- Request ID. The
x-request-idheader is copied into the result and into every exception, ready for your logs.
Using the client
Here is the client triaging a support message. The questions are an ordinary dict, so you can keep them in one module and reuse them across calls.
from jev_client import JevClient
QUESTIONS = {
"department": {
"type": "choice",
"instructions": "Which team should handle this message?",
"criteria": {
"billing": "Charges, invoices, refunds or subscription changes",
"technical": "Bugs, outages, errors or integration problems",
"sales": "Pricing questions or new purchases",
},
},
"is_urgent": {
"type": "noul",
"instructions": "Does this message need a reply today?",
},
}
client = JevClient(read_timeout=20)
result = client.decide(
state="Our checkout has returned errors since this morning and no orders are coming in.",
questions=QUESTIONS,
)
department = result["answers"]["department"]
urgent = result["answers"]["is_urgent"]["noul"]
print(department["choice"], round(department["confidence"], 2), round(urgent, 2))
print(result["model"], result["usage"]["input_tokens"], result["request_id"])
Processing a batch without burning credit
Batch jobs are where error handling matters most. When the balance runs out halfway through 10,000 tickets, every remaining call will return 402. The job should stop at the first one, not hammer the API 9,000 more times. The same applies to a revoked key.
import logging
from jev_client import JevAuthError, JevBalanceError, JevClient, JevError
log = logging.getLogger("triage")
def triage_all(tickets, client: JevClient, questions: dict):
results = {}
for ticket in tickets:
try:
out = client.decide(state=ticket["text"], questions=questions)
except (JevBalanceError, JevAuthError) as exc:
log.error("Stopping batch: %s (request_id=%s)", exc, exc.request_id)
break
except JevError as exc:
log.warning("Skipping ticket %s: %s (request_id=%s)", ticket["id"], exc, exc.request_id)
continue
results[ticket["id"]] = out["answers"]
log.info("ticket=%s model=%s input_tokens=%s request_id=%s", ticket["id"],
out["model"], out["usage"]["input_tokens"], out["request_id"])
return results
If you add concurrency, for example with concurrent.futures.ThreadPoolExecutor, keep the number of workers small and grow it slowly. A rising count of 429 responses means you are going faster than the model service accepts. Jev API does not publish a fixed rate limit, so treat 429s as the signal and let the backoff absorb short bursts. Note that requests.Session is not officially thread-safe, so give each worker thread its own JevClient.
Testing the client without spending credit
Retry and error paths are hard to trigger against the live API on demand, so test them with fake responses. unittest.mock can replace session.post. These tests check that a 402 is not retried and that a 503 is:
import json
from unittest import mock
import pytest
import requests
from jev_client import JevBalanceError, JevClient
QUESTION = {"q": {"type": "noul", "instructions": "Is this spam?"}}
def fake_response(status, body, headers=None):
resp = requests.Response()
resp.status_code = status
resp._content = json.dumps(body).encode()
resp.headers.update(headers or {})
return resp
def test_402_is_not_retried():
client = JevClient("jev_test_xxx", max_retries=3)
body = {"error": {"code": "insufficient_balance", "message": "Balance too low", "request_id": "req_1"}}
with mock.patch.object(client.session, "post", return_value=fake_response(402, body)) as post:
with pytest.raises(JevBalanceError) as info:
client.decide("text", QUESTION)
assert post.call_count == 1
assert info.value.request_id == "req_1"
def test_503_is_retried_then_succeeds():
client = JevClient("jev_test_xxx", max_retries=3, backoff_base=0)
unavailable = fake_response(503, {"error": {"code": "service_error", "message": "x", "request_id": "req_2"}})
ok = fake_response(200, {"model": "jev-1.13.0",
"answers": {"q": {"type": "noul", "noul": 0.1}},
"usage": {"input_tokens": 30, "output_tokens": 5}})
with mock.patch.object(client.session, "post", side_effect=[unavailable, ok]) as post:
result = client.decide("text", QUESTION)
assert post.call_count == 2
assert result["answers"]["q"]["noul"] == 0.1
Logging and secrets
Log the response model, usage.input_tokens and the request ID for every call. The model tells you which version produced each decision. The token count lets you reconcile your own records with the usage table on your dashboard. Never log the Authorization header or the key itself. If you dump request objects while debugging, strip the header first. If a key does end up somewhere it should not, revoke it on the dashboard and create a new one. Revocation takes effect immediately for new requests.
Next steps
With a client that fails safely, the next improvements are better questions and a cost budget:
- Docs: account, key and credit setup from start to finish.
- API reference: the full request and response contract for
POST /v1/systemoneandGET /v1/models. - Pricing: current per-token price and how prepaid credit works.
More tutorials on this blog:
- Getting started with Jev API: your first structured decision in 5 minutes
- Using Jev API from Node.js and TypeScript with typed, validated responses
- Designing decision questions: how to frame state and criteria for reliable typed answers
- Controlling cost on Jev API: estimating tokens, tracking usage and managing prepaid credit