Jev API does one thing: it evaluates a state against typed questions and returns structured answers with probabilities. The model does the judging, but you define what is being judged. Vague options, overlapping categories and states padded with noise all produce answers that are hard to act on. This guide covers how to frame decisions so that the answers are stable, interpretable and easy to wire into code.
Everything below works within what the request actually accepts. A request has model, state and questions. Each question has a type (choice, score or noul), instructions and, for choice and score, criteria. There are no hidden parameters to tune. The quality of the decision comes from those fields.
Start from the action, not the prompt
Before writing a question, write down what your code will do with each possible answer. If the answer is “billing”, the ticket goes to the billing queue. If urgency is at the top level, someone gets paged. If you cannot name an action for every outcome, the question is not ready yet. Either some options do not matter, or the decision your system needs is a different one.
This habit also tells you when not to use Jev at all. If the outcome follows from a rule you can state exactly (“orders over $10,000 need manager approval”, “accounts younger than a day cannot change their payout address”), put that rule in code. It is cheaper, deterministic and auditable. Use Jev for the judgment calls that rules handle badly: tone, intent, topic, relevance and urgency.
Pick the question type that matches the decision
choice: exactly one of several options
Use choice when exactly one option applies, such as a routing destination, a category or a next step. criteria is an object. The keys become the possible values of choice in the answer, so make them stable, machine-friendly identifiers. The values define what each option means.
"department": {
"type": "choice",
"instructions": "Which team should own this request?",
"criteria": {
"billing": "Payment, subscription, or invoice issue",
"technical": "Bug, outage, or integration failure",
"sales": "Pricing or account inquiry"
}
}
score: a position on an ordered scale
Use score when the options have a natural order, such as urgency, severity, quality or fit. criteria is an array, ordered from the lowest level to the highest. The answer returns a legend mapping "0", "1", … back to your text, a probability for each level, and a numeric score. Because the scale is ordered, “between level 1 and level 2” is a meaningful answer here in a way it is not for choice.
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm and factual", "Frustrated but polite", "Very angry or confrontational"]
}
noul: a degree of yes
Use noul for a yes/no judgment where the strength of the yes is useful: is this spam, does this message ask for a refund, should this be escalated. The answer is a single noul value from 0 to 1. Instructions alone are enough. As the homepage example shows, you can also describe what “true” and “false” mean in criteria when the boundary needs spelling out.
"is_urgent": {
"type": "noul",
"instructions": "Does this message convey urgency?",
"criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" }
}
A quick test: if you would draw it as radio buttons, use choice. If you would draw it as a slider with labelled stops, use score. If you would draw it as a single checkbox and care how sure the answer is, use noul.
Write the state as evidence
The state is the evidence. It can be a string, an object or an array, so send whatever represents the facts most clearly.
- Include what the decision depends on, and nothing else. For ticket routing, the subject and body matter. The full HTML email with signatures, legal footers and quoted history mostly adds tokens and noise.
- Use an object for records. Named fields make it clear which fact is which:
"state": {
"subject": "Charged twice this month",
"body": "I see two identical charges on my card for the March invoice.",
"customer_plan": "Team",
"previous_tickets_last_30_days": 0
}
- Do not put instructions in the state. “Please classify this as billing” inside the state blurs the line between evidence and question. Instructions belong in
instructions. - Do not add facts you do not have. If the customer's plan is unknown, leave it out rather than guessing. The model can only weigh what you send.
- Mind the size. The model pages list a 64k total context, with the state plus the longest question limited to 32k. Size matters for cost too, because you pay per input token.
Write criteria that can be told apart
Most weak decisions come down to criteria that overlap or are left undefined. A few rules fix most of them.
Put the meaning in the criteria
An option called technical with the description “Technical” forces the model to guess where your boundary lies. “Bug, outage, or integration failure” tells it. Keep instructions short and phrased as a question, and do the defining in criteria.
Make options mutually distinguishable
Compare these two option sets for the same ticket queue:
// Overlapping: a failed payment is both "billing" and "problem"
{ "billing": "Money related", "problem": "Something is broken", "other": "Anything else" }
// Distinguishable: each case has one natural home
{
"billing": "Charges, invoices, refunds, or plan changes, where nothing is malfunctioning",
"technical": "Something in the product is failing or behaving incorrectly, including payment errors",
"account": "Login, access, or user management",
"other": "None of the above"
}
The second set says where the borderline case (a failing payment form) belongs. When you find a case that fits two options in your own data, edit the descriptions until it fits one.
Close the option set
choice always returns one of your keys. If real inputs can fall outside every option, add an explicit other or none option. Otherwise the model has to force them into the nearest wrong bucket. Routing other to a person is usually the right default.
Anchor score levels with observable descriptions
“Low / Medium / High” leaves each level open to interpretation. “Routine / Needs attention this week / Needs attention today / Production is down” describes something you could check. Keep the levels in strict order from lowest to highest, since the response legend relies on that order.
Ask one thing per question
“Is the customer angry and likely to cancel?” is two questions. When the answer is 0.6, you cannot tell which half drove it. Split it into is_angry and cancellation_risk.
Multi-label and multi-dimension decisions
A message can mention billing and a bug. choice picks exactly one option, so do not use it for labels that can apply together. Ask a separate noul question for each label instead, such as mentions_billing and mentions_bug, and apply your threshold to each one.
When several questions concern the same state, send them together in one request. Each question needs a stable key (department, urgency, should_escalate), and the answers come back under the same keys. One request keeps the state in one place and gives you one response to log. If token cost matters to you, compare usage.input_tokens for a combined request against separate ones on your own payloads.
Turn probabilities into actions
The API returns probabilities and confidence values. It does not decide where your cut-offs are. A noul of 0.62 is not automatically a yes, and 0.5 is not automatically the right threshold. A practical way to set thresholds:
- Collect a few dozen to a few hundred real inputs and label them yourself with the answer you would want.
- Run them through your questions and record the returned
noul,choice,confidenceand probabilities. - For each threshold you are considering, count how many cases would be handled wrongly in each direction. Choose based on which mistake costs more. Missing an urgent ticket is usually worse than paging someone unnecessarily.
- Add a review band. Below a confidence level you choose, or when the top two
choiceprobabilities are close, send the case to a person instead of acting automatically.
Keep the labelled set. Whenever you change the criteria or the model, rerun it and compare before deploying.
Pin the model once thresholds matter
jev-latest follows the current release and jev-preview follows the newest one, so the probability distribution behind an alias can shift when a new release ships. While you are exploring, that does not matter. Once you have tuned thresholds against a labelled set, switch to the pinned jev-1.13.0 so they stay valid. Moving to a newer release then becomes a deliberate step: rerun the labelled set, adjust thresholds, deploy. The response model field reports the version that served each request, so log it.
A checklist before you ship
- Every possible answer maps to a concrete action in your code.
- Rules that can be evaluated exactly live in code, not in a question.
- The state contains the relevant facts only, with no instructions or invented details.
- Choice options are defined in
criteria, do not overlap, and includeotherwhere needed. - Score levels are ordered lowest to highest and describe observable conditions.
- Each question asks one thing. Labels that can co-occur are separate
noulquestions. - Thresholds come from a labelled set, with a human review band for uncertain cases.
- The model is pinned once thresholds are tuned.
Next steps
Once your questions are solid, wire them into a client and keep an eye on spend:
- Docs: account, key and credit setup from start to finish.
- API reference: the full request and response contract for
POST /v1/systemoneandGET /v1/models. - Pricing: current per-token price and how prepaid credit works.
More tutorials on this blog:
- Getting started with Jev API: your first structured decision in 5 minutes
- Calling Jev API from Python: a robust client with timeouts, retries and error handling
- Using Jev API from Node.js and TypeScript with typed, validated responses
- Controlling cost on Jev API: estimating tokens, tracking usage and managing prepaid credit