Jev API is billed by input tokens, from a prepaid balance. Output tokens are free, and there are no seats or subscriptions. That makes costs easy to predict, but only if you know how a charge is calculated, what happens when the balance runs low, and where to see what you have spent. This guide covers all three, with code to estimate and track spend from your own application.
How a call is billed
The price is $0.084 per 1 million input tokens. Output tokens cost $0. The cost of a call is based on the usage.input_tokens returned in the response:
cost (USD) = input_tokens × 0.084 / 1,000,000, rounded up to the next $0.000001
Balances and charges are tracked in micro-dollars (millionths of a dollar), and each call is rounded up to a whole micro-dollar. The example response on the API reference reports 392 input tokens: 392 × 0.084 = 32.928 micro-dollars, which rounds up to $0.000033. The 65 output tokens in the same response add nothing.
Credit is deducted only after a call succeeds. A request that ends in an error is not charged. That includes 400, 401, 402, 429 and 5xx responses. All three model IDs (jev-latest, jev-preview and jev-1.13.0) are listed at the same price, so choosing between them is a quality and stability decision, not a cost one.
The balance check and HTTP 402
Before a request reaches the model, Jev API reserves a conservative estimate of its cost and checks that your balance covers it. The estimate assumes each question may be evaluated together with the whole state, so the state is counted once per question. Lengths are measured on compact JSON and divided by 4, rounding up:
reserve tokens = 256 + sum over questions of (state tokens + question tokens + 64)
The 256 and 64 are fixed allowances for the framing added around the request and around each question. The reserve is priced like any other input tokens. If your balance is zero or negative, or lower than the reserve, the call fails immediately:
{
"error": {
"code": "insufficient_balance",
"message": "Balance 0.000012 USD is below the estimated 0.000056 USD charge.",
"request_id": "req_…"
}
}
The status is 402, and nothing is charged. Three consequences for your code:
- Never retry a 402 automatically. It will keep failing until someone adds credit. Stop the job and alert a person, as shown in the Python client guide.
- The reserve is not the bill. The charge uses the token count reported after the call. The reserve is deliberately generous, so most calls cost less than it.
- A balance can dip slightly below zero. If a call reports more tokens than were reserved, the full amount is still charged. Every later call is then rejected with 402 until you add credit.
For the three-question request on the API reference, the reserve is 662 tokens (56 micro-dollars), and the example response reports 392 input tokens (33 micro-dollars). Use the reserve to understand the 402 check. Budget from the real usage.input_tokens numbers.
Estimating tokens before you send
You can reproduce the reserve locally. This is useful for spotting oversized payloads before they go out, and for working out the smallest balance a given request needs. In JavaScript or TypeScript:
const tokens = (value: unknown) => Math.ceil(JSON.stringify(value).length / 4);
export function estimateReserveTokens(state: unknown, questions: Record<string, unknown>): number {
const stateTokens = tokens(state);
let total = 256;
for (const [key, question] of Object.entries(questions)) {
total += stateTokens + tokens({ [key]: question }) + 64;
}
return total;
}
export function costMicroUsd(inputTokens: number): number {
return Math.ceil((inputTokens / 1_000_000) * 0.084 * 1_000_000);
}
And in Python. The details that matter are compact separators and ensure_ascii=False, which make the string match JavaScript's JSON.stringify:
import json
import math
def _tokens(value) -> int:
return math.ceil(len(json.dumps(value, separators=(",", ":"), ensure_ascii=False)) / 4)
def estimate_reserve_tokens(state, questions: dict) -> int:
state_tokens = _tokens(state)
return 256 + sum(state_tokens + _tokens({key: q}) + 64 for key, q in questions.items())
def cost_micro_usd(input_tokens: int) -> int:
return math.ceil((input_tokens / 1_000_000) * 0.084 * 1_000_000)
Two caveats. First, this is a heuristic, not a tokenizer. The authoritative number is the usage.input_tokens in each response, so measure real calls before committing to a budget. Second, JavaScript counts string length in UTF-16 code units while Python counts code points, so text with emoji or other characters outside the Basic Multilingual Plane gives slightly different lengths. For a budget estimate, that difference rarely matters.
The cost_micro_usd functions use the same arithmetic and rounding as the billing formula above. Run a few representative requests, note the real input_tokens, and use the average for planning.
Back-of-envelope budgeting
Suppose your typical request uses 400 input tokens. That is 400 × 0.084 = 33.6 micro-dollars, billed as 34 per call:
| Calls | Cost at 34 micro-dollars per call |
|---|---|
| 1,000 | $0.034 |
| 10,000 | $0.34 |
| 100,000 | $3.40 |
| 1,000,000 | $34.00 |
A worked example: the support-triage request in the getting started guide reserves 671 tokens (57 micro-dollars) before it runs. Suppose your logs show it really uses about 310 input tokens per call. Then each call costs 27 micro-dollars, and 3,000 tickets a month cost 81,000 micro-dollars, or about $0.08. Measure, multiply, and add headroom for the long messages that inevitably turn up.
At 400 tokens per call, the minimum $5 top-up covers about 147,000 calls. The arithmetic is simple, but run it with your measured token count. A state containing a long email thread can easily be ten times larger than a one-line message.
Monitoring usage and balance on the dashboard
The dashboard shows three things relevant to cost:
- Balance, shown to six decimal places, at the top of the page.
- Usage: your most recent 20 calls to
/v1/systemone, each with its time, model, input tokens and cost. - Ledger: your most recent 20 balance movements. Top-ups appear as
topupand call charges asusage, with a note such asjev-latest 392 in / 65 out.
The dashboard is good for spot checks: confirming a top-up arrived, or checking whether a call that timed out on your side was actually charged. It shows only recent rows and does not break usage down by API key. For anything longer-term, keep your own records.
Tracking spend in your own code
Every successful response includes usage.input_tokens, and every response includes an x-request-id header. Record both, along with the response model and which feature or customer the call was for. Then you can answer questions the dashboard cannot, such as which feature costs the most or how spend grew this month. A minimal budget guard in Python:
import math
import threading
class Budget:
"""Stop making calls once a local spending limit is reached."""
def __init__(self, limit_usd: float):
self.limit_micro = int(limit_usd * 1_000_000)
self.spent_micro = 0
self._lock = threading.Lock()
def record(self, input_tokens: int) -> None:
cost = math.ceil((input_tokens / 1_000_000) * 0.084 * 1_000_000)
with self._lock:
self.spent_micro += cost
def exhausted(self) -> bool:
return self.spent_micro >= self.limit_micro
daily = Budget(limit_usd=1.00)
# in your request loop:
# if daily.exhausted(): stop and alert
# result = client.decide(...)
# daily.record(result["usage"]["input_tokens"])
A local budget is not a substitute for the server-side balance. It is a guard you control, so a bug that loops on the API burns through a dollar instead of your whole balance.
Reducing tokens per call
Since only input is billed, every saving comes from sending less:
- Trim the state. Strip HTML, email signatures, quoted replies and boilerplate. Send the fields the decision depends on, not the whole database row.
- Keep criteria concise. Clear, specific descriptions help accuracy. Long, repetitive ones mostly add tokens. The question design guide covers how to keep them precise.
- Group questions about the same state into one request, then compare
usage.input_tokensagainst separate calls on your own data. - Do not ask what code can answer. Exact rules, lookups and arithmetic are free in your own code.
- Avoid duplicate work. If the same input can arrive more than once, for example from re-delivered webhooks or re-run jobs, store the decision keyed by a hash of the payload and reuse it.
- Be careful with retries. Retrying error responses costs nothing, because failed calls are not charged. Retrying after a client-side timeout can pay for the same decision twice.
Best practices for prepaid credit
- Top up in measured amounts. The minimum is $5, entered in whole dollars on the dashboard and paid through Stripe Checkout. Purchased credit is non-refundable and cannot be withdrawn (see the Terms of Service), so size top-ups to measured usage rather than guesses.
- Allow for confirmation time. Credit is applied when Stripe's signed webhook confirms the payment. Top up before the balance hits zero, not after jobs start failing with 402.
- Remember that test keys are billed too.
jev_test_andjev_live_keys draw from the same balance. Development traffic costs the same as production traffic. - Use one key per service and revoke unused ones. Any request made with an active key is billed to your balance. Revoking a key on the dashboard stops new requests with it immediately.
- Alert on 402. Treat the first
insufficient_balanceresponse as an operational alert, the same way you would treat a full disk.
Next steps
Costs are predictable once you measure tokens per call. To go further:
- Docs: account, key and credit setup from start to finish.
- API reference: the full request and response contract for
POST /v1/systemoneandGET /v1/models. - Pricing: current per-token price and how prepaid credit works.
More tutorials on this blog:
- Getting started with Jev API: your first structured decision in 5 minutes
- Calling Jev API from Python: a robust client with timeouts, retries and error handling
- Using Jev API from Node.js and TypeScript with typed, validated responses
- Designing decision questions: how to frame state and criteria for reliable typed answers