Jev: A Decision Model That Classifies Without Generating Text
Madhu Shantan
Sep 25, 2026 · 6 min read

Why AI Classification Needs a Decision Model, Not an LLM
Many AI workflows and usecases need a small decision rather than a written answer. A support ticket needs a queue, an invoice needs a category to be put in, a retrieved passage needs a relevance check, and a user request needs the right model to route to. Traditional ML classifiers handle these jobs well when the inputs and labels are stable. When the context is messy or the criteria change, teams often reach for an smaller, faster LLM - and that’s the default option today. But it adds latency (~2sec atleast): the model generates a structured answer token by token before the application can parse it and act.
Jev takes a different approach. TypeSafe built it to accept context and typed questions, then return probabilities over allowed answers directly. In TypeSafe’s published workflow comparisons, Jev was 193.6 times faster and 444.6 times cheaper than the LLM alternatives. Those gains will depend on the workload, but the underlying idea is useful regardless of the exact numbers. If the output your usecase/workflow needs is a decision, why make the model write one first?
Typed Questions vs. Structured Outputs: How the Jev API Works
Imagine a customer writes: “I’ve been charged twice, nobody has replied, and I need this fixed today.” You need to identify the right team, judge urgency, and decide whether the charge might be unauthorized. A chat model could put those answers in JSON, but a confidence number written into that JSON is not automatically a calibrated probability.
Jev’s request shape makes the decision explicit. The state holds the customer message, perhaps alongside account history and refund policy. The questions map asks for several independent judgments about that same state. There are three types of primitives/questions jev can answer in :
- Choice : selects one option from a defined set, such as a ticket queue, and returns a probability for each option.
- Score : rates something against ordered descriptions, such as levels of urgency, and returns a score and probabilities across those levels.
- Noul : evaluates a yes-or-no question and returns the probability that the answer is yes, from zero to one.
One request can ask all three questions about the same customer message:

The question types reflect the decisions your code needs. A queue is a closed set of alternatives, so Choice fits. Urgency has an order, so Score can place a case between adjacent levels. Whether the customer alleges an unauthorized charge is one proposition, so Noul fits. You define the answer space, then let code decide how to use it.
This changes how you compose a workflow. Ask narrow questions, then combine their outputs in ordinary code. The model handles judgments that are hard to write as rules. Your code owns the rules themselves. If your refund policy changes, you change the decision logic rather than hoping a rewritten prompt produces the right behavior.
Why Jev Is Faster and Cheaper Than LLM Token Generation
Chat models generate output one token at a time. Even for a three-way classification, the model has to spell out JSON keys, a label, and perhaps a confidence field. Constrained formats can keep that text valid, but they don’t eliminate sequential decoding. TypeSafe says Jev returns answer probabilities in parallel. The server serializes the numbers into JSON after the model has made its decisions.
The other important optimization is shared work. Suppose you have a long incident report and 30 questions about it. Separate classification calls process that report 30 times. Jev accepts one state with many questions; each sees the same state and is evaluated independently. The state counts once against the request’s token budget, and Jev’s pricing currently charges only for input tokens. At the listed $0.042 per million input tokens, a request billed for 1,000 input tokens costs about $0.000042. The exact bill depends on the state and questions you send, but the API makes it very economical to ask several questions about the same evidence.
What happens inside the model is less certain. Archer Hume’s black-box investigation found behavior consistent with processing the state once and handling question branches together. Information placed in the shared state affected another question, while information placed in a sibling question did not. Latency stayed nearly flat as the number of questions grew to around 100 before rising more steadily. A shared prefix with isolated branches is plausible. A particular KV-cache layout or attention mask remains an inference.
The same caution applies to the output layer. A straightforward design would turn a hidden representation into logits for the allowed answers, then apply a softmax. Jev may use a dedicated decision head, reserved label tokens, or another readout with the same visible behavior. Hume argues for a causal transformer and considers sparse experts likely, but TypeSafe has not disclosed enough to establish either. What developers can depend on is the typed numerical output, not a particular internal design.
Confidence Score vs. Probability: Why Model Calibration Matters
A Choice answer gives you a distribution over the options. Suppose the probabilities for billing, support, and other are 0.81, 0.12, and 0.07. Your software sees the leading answer and the uncertainty around it. You can route clear cases automatically and send ambiguous ones for review.
But a probability in an API response isn’t automatically trustworthy. Calibration means that, across cases assigned roughly 0.8 probability, the event should occur about 80% of the time. It says nothing certain about one case. TypeSafe calls its post-training approach Reinforcement Learning for Calibrated Decisions, or RLCD. The exact training recipe is not public, and calibration still needs to be checked on your own traffic.
There is a subtle API detail here: confidence is not a separate learned prediction that an answer is correct. For Choice, TypeSafe derives it from the probability distribution. With three options and a top probability of 0.8, the confidence score is 0.7 after normalizing its distance from a uniform three-way guess. That summarizes how peaked the distribution is. It is not a substitute for the probability of the outcome you care about.
Think about the unauthorized-charge question. Missing a real case may cost far more than reviewing an innocent one, so you might send it for review even when noul is only 0.15. The threshold belongs to your policy. Jev supplies the signal; your code decides what the risk justifies.
Jev Limitations: Where Decision Models Break in Practice
A typed output guarantees the answer has the expected shape. It doesn’t guarantee the answer is right. TypeSafe’s guide for Jev 1.13 says the model can struggle with exact arithmetic, counting, date comparisons, indirection, irrelevant context, and adversarial content. It also cannot write the customer reply after routing the ticket. Jev fits the judgment in the middle of a workflow; code and other models handle the work around it.
Even a well-formed Choice can be sensitive to how you frame the options. Hume found that reversing option order moved a support classification’s leading probability from roughly 0.85 to roughly 0.95. Adding an irrelevant fifth option changed the relative odds between two existing choices in repeated probes. Neither result reveals where the interaction happens inside Jev, but both matter to anyone thresholding the output. If 0.9 is your automatic-routing cutoff, an option-order change could move the same case across it.
So ideally we should evaluate Jev on real examples from the target workflow. Test alternate option orders and missing-context cases. Measure calibration per important class, alongside accuracy and the cost of each kind of mistake. Pin a model version if your thresholds depend on its behavior, because the jev-latest alias can move. Measure end-to-end latency, including network time and fallback. TypeSafe’s published comparisons are compelling, but your workflow is the test that counts.
Conclusion
Many applications need to turn messy context into a bounded judgment, expose uncertainty, and decide what happens next. Jev makes that a model interface. Its speed and price make frequent classification practical, while its probability distributions let you build explicit routing and review policies. The hard part is defining good questions, validating the probabilities on your data, and deciding which mistakes your system can tolerate.
Together AI also recently released Tev1-4B-experimental, a Jev-inspired classifier fine-tuned from Qwen3.5-4B, shows that this idea is already spreading. I’m excited to see where these decision models fit: for the right tasks, they could deliver useful judgments far faster and cheaper than a general-purpose LLM.
That’s why Jev is relevant to us at Bifrost. Jev is now available as a first-class provider, so applications can send decision requests through the gateway. We’re also working on an opt-in Jev option for the Complexity Router, where a decision model can help choose the right model before the expensive call happens. Use generative models when you need generation, and a decision model when the job is to decide.
For the deeper architectural investigation and the evidence behind it, read Jev’s Architecture Unmasked. The TypeSafe documentation covers the API contract and current model limits.