← All posts

Jev and structured outputs: when AI should choose, not write

Most of what a business asks AI to do is not writing. It is deciding: which queue a ticket belongs in, whether an invoice matches its order, whether a message is angry, whether a request is safe to run. For decisions like these you want the model to pick one of your options, with a probability, not compose a paragraph. Structured outputs give you the format. Decision models such as Jev, which TypeSafe AI released on 15 September 2026, go one step further and return only a choice and a probability. Used well, both are faster, cheaper and easier to check than free text. Used badly, they give you a wrong answer in a very clean shape.

What is a typed decision?

A typed decision is an AI answer restricted to a set of options you define in advance, such as refund, replace or escalate, or a yes or no, or a score from 1 to 5. The model cannot answer outside that set, so the output can be stored, counted and audited like any other field in a database. Classifying a support email, routing a lead, flagging a risky payment and tagging a document are all typed decisions. In most business systems I build, the typed decisions outnumber the places where the model genuinely has to write.

What do structured outputs guarantee, and what do they not?

Structured outputs guarantee the shape of an answer, not its correctness. OpenAI, Anthropic and Google all enforce a JSON schema by grammar-constrained sampling, so the reply always parses. The value inside the valid JSON can still be wrong, an enum can still drift, and a long answer can still be cut off. Two practical rules follow. Keep a validate-and-repair step even when the schema is enforced. And let the model reason first and fill the structure last, because constraining the format measurably costs accuracy on hard reasoning tasks (Tam et al., 2024).

The provider documentation is the reference here: OpenAI structured outputs, Anthropic structured outputs and Gemini structured output.

What is Jev?

Jev is a "System One" decision model from TypeSafe AI, released on 15 September 2026. You send it a state, as text or JSON, and a set of typed questions. It returns, for every question, one of the options you defined plus a probability for each option. According to TypeSafe, a call takes about a tenth of a second and costs USD 0.042 per million input tokens, with output free. Jev does not chat and does not write. Those speed and cost figures are the vendor's own, measured from the US West Coast, and should be treated as claims until you measure them on your own traffic.

What did the independent tests find in the first week?

The independent tests found that Jev is fast and cheap, sits roughly level with mid-priced language models on quality, and depends heavily on how the question is asked. Three results stand out:

  • One broad question loses, five narrow ones win. On 2,000 phishing and legitimate emails, asking Jev one question ("is this phishing?") scored 62.6% accuracy. Asking five narrow questions in the same call and combining the answers with a small logistic regression scored 95.0% (jev-phishing-bench, 17 September 2026).
  • Option names change decisions. A September 2026 paper found that binding different names to the same options, for example "0" and "1" instead of "no" and "yes", changed the hosted model's decisions enough to drop its AUC from 0.81 to 0.58 (arXiv 2609.26758).
  • Confidence is not the same as knowing. On a question whose answer depended on a hidden policy, so no model could know it, Jev still answered at an average probability of 0.74 (jev-ood-calibration).

Accuracy also fell on languages other than English in the early reports. For a Greek business that is the line to test first, not the headline speed.

How should you design a decision task for AI?

Design the decision so the model answers small, checkable questions and the code does everything else. The rules below hold for Jev, for structured outputs on any large model and for a classifier you train yourself:

  1. Split broad questions into narrow ones. "Does the email contain a link whose text and target differ?" beats "is this suspicious?". Combine the narrow answers with a rule or a small fitted model.
  2. Keep arithmetic, dates and totals in code. A decision model is weak at numbers and counting. Compute the VAT difference yourself and ask the model only whether the explanation on the invoice is plausible.
  3. Always offer an "unknown" option. A model forced to choose between two wrong answers will pick one with confidence.
  4. Use neutral, descriptive option names and keep them fixed once you have calibrated.
  5. Calibrate on your own labelled examples. Fifty to three hundred labelled cases are enough to learn where the probability can be trusted on your data.
  6. Route low-confidence cases to a larger model or to a person, and log every routed case.

What does letting a model say "I don't know" cost?

Letting a model decline costs a little coverage and buys a lot of trust. I measured this on my own small model, a 5-million-parameter intent classifier that runs in the browser and that you can try at watch it think. On 10,578 held-out sentences it picks the right intent 74.53% of the time when it must always answer. Allowed to decline below 0.50 confidence, it scores 72.63%, stays silent on 834 sentences and hands those to a fallback instead of guessing. Every decision system I ship has that branch, because a wrong action costs more than a question.

Where do typed decisions pay off for a Greek business?

Typed decisions pay off wherever the same judgement is made hundreds of times a day. Typical cases: routing incoming emails in Greek and English to the right person, tagging supplier invoices before bookkeeping, marking which e-shop orders need a manual check, and screening form submissions for spam. Each one is a small, repeatable choice with a cost when it is wrong, which is exactly what you can measure before going live. Test on your own Greek messages, because published results in other languages will not transfer one to one.

These patterns sit inside the agent systems I run in production, and the question of when a plain script is still the better answer is in when not to build an AI agent.

Frequently asked questions

Is Jev a large language model?

No. Jev is marketed as a "System One" model that returns typed decisions with probabilities instead of text. It cannot hold a conversation or write a reply, so it complements a language model rather than replacing one.

Do structured outputs stop hallucinations?

No. Structured outputs guarantee valid format, not correct content. A model can still choose the wrong option or invent a value inside a perfectly valid JSON object, so validation and an "unknown" option are still required.

Can I use decision models on Greek text?

Yes, but measure first. Early independent reports showed accuracy dropping on non-English text, so build a labelled Greek test set from your own data before relying on any number.

If you have a decision your team makes a hundred times a day, send me the decision and what a wrong answer costs you. That is enough to tell whether it should be a rule, a classifier or an agent.