# How Jev Works Under the Hood: An Illustrated Guide

By Aish Reganti (Aishwarya Naresh Reganti) · Published September 22, 2026

Canonical article: https://aishreganti.com/blog/how-jev-works-under-the-hood/

![A generative LLM writes tokens in sequence; a decision model returns probabilities for supplied choices. The percentages are an example.](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/hero.png)

Jev is a new model from TypeSafe AI that’s been doing the rounds on social media, with people talking about how fast it is and what they’re building with it. TypeSafe describes it as the first release in a class it calls **System One models**.

The System One name borrows from the idea of fast, intuitive judgments. For Jev, that means taking some context and questions with defined answers, then returning decisions and probabilities. I’ll use the word *decision model* to describe this kind of model. <sup>[1](#ref-1)</sup>

For example, we could give Jev a customer’s message and 4 departments, then ask which department should handle it. It would return a department from that list, along with probabilities for the choices. A generative LLM could answer the same question, but it would produce text, even if that text were only a department name or a JSON response.

Jev keeps its output within the choices or values we define. Depending on the question, those outputs take 1 of 3 forms:

| Output | What you ask | What comes back |
| --- | --- | --- |
| Choice | Which department should handle this? | 1 of your supplied departments, probabilities for each option, and confidence. |
| Score | How urgent is this, using these defined levels? | A numerical rating based on your ordered levels, probabilities for the levels, and confidence. |
| Noul | Does this customer request a refund? | A number from 0 to 1 representing the probability that the answer is yes. |

Those names come from [Jev’s question types](https://docs.typesafe.ai/primitives). For Score, you define what the levels mean. For Noul, 0.8 means an 80% probability of yes.

I was super curious to go a little deeper into the weeds and understand how these models work and how they differ from traditional LLMs. I put together this visual guide to make that easier to follow.

Heads up: since Jev itself isn’t open source, I drew on open-source implementations like [Open-Jev](https://github.com/Zefan-Cai/Open-Jev) and [Laya](https://github.com/NandhaKishorM/laya) for many of the details below. They help us understand how decision models can be built, though Jev’s exact design isn’t public.

## 1. How a generative LLM produces text

Before getting into decision models, it helps to refresh how an LLM generates text. We can then follow what changes when the model needs to choose an answer from a supplied set of options.

Let’s use 1 customer message throughout: “My card was charged twice.” First, we ask a generative LLM to draft a reply. It might write, “I can help check the duplicate charge.” We’ll follow how it produces that reply through 3 components, then keep the same numbers when we adapt the model to choose a support department.

### Component 1. Tokenization and embeddings

The customer’s message starts as text, which we need to turn into numerical inputs for the model. A tokenizer breaks it into pieces called tokens. For our example, imagine `My`, ` card`, ` was`, ` charged`, ` twice`, and `.`. Tokens can also be word fragments or punctuation, and each gets an ID that the model uses to look up an *embedding*, a list of numbers representing that token.

The tokenizer prepares the input outside the neural network. The embedding layer is inside it. Together, they turn the text into something the model can process.

![Tokenization and embeddings](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/inputs.png)

[Watch this animation](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/inputs.mp4)

### Component 2. The language-processing layers

Those embeddings give us an initial representation of each token. As they pass through a stack of layers, often called the **backbone**, attention lets each position use information from other positions in the available context. A feed-forward network then processes each representation further. The model also accounts for position, because word order matters.

Think of “charged” in “my card was charged twice.” Its surrounding words help establish that we’re talking about a payment. Attention helps the model combine that context as it builds a representation of the message.

![Inside the language-processing layers](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/layers.png)

[Watch this animation](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/layers.mp4)

Which surrounding words attention can use depends on how these layers are arranged. 2 arrangements we’ll need to understand are the **encoder** and the **decoder**, often described as handling understanding and generation respectively.

Think about listening to someone and then responding. While listening, you’re trying to make sense of what they’re saying. When you respond, you turn what you understand into words, and what you’ve already said shapes what comes next. That’s roughly how the encoder and decoder divide the work in an encoder-decoder model.

An encoder can use context from both directions in the supplied text. A decoder uses the current and preceding positions, which fits continuing a sequence. In an encoder-decoder model, the encoder represents the input and the decoder uses that representation to generate the output. <sup>[2](#ref-2)</sup>, <sup>[3](#ref-3)</sup>

But many generative LLMs, including Qwen’s language models, are decoder-only. Their decoder layers also process and interpret the input; they don’t need a separate encoder to do that. So “understanding side” and “generation side” describe their usual roles, rather than 2 halves that every LLM contains. <sup>[4](#ref-4)</sup>

![Encoder and decoder context](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/context.png)

### Component 3. The output head

After those decoder layers have processed the message, we still have numerical representations. To begin writing the reply, the **language-model head** takes the final position’s representation and scores the possible next tokens across the vocabulary. Those scores become probabilities, and a token is selected.

Suppose the first token of the reply is `I`. It joins the sequence after the prompt, its embedding goes through component 2, and component 3 predicts again. The next token might be ` can`, followed by ` help`, as the model writes “I can help check the duplicate charge.”

Each new token becomes context for the next token. This is the answer-generation loop, also called autoregressive generation. The animation follows the first 3 tokens; the same process continues through the rest of the reply.

During this loop, the model reuses stored information about earlier tokens, so it doesn’t start from scratch each time. It still has to do more computation for every new output token. <sup>[5](#ref-5)</sup>

![Predict a token, then repeat](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/generation.png)

[Watch this animation](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/generation.mp4)

We needed that loop to write a reply. Now suppose the application only needs to route the same message: “My card was charged twice.” We supply 4 departments: billing, fraud, card replacement and general support. The model still needs to interpret the message, but the answer is limited to those 4 options. An LLM can generate a department name; a decision model can score the options directly and return their probabilities.

## 2. 2 ways to build a decision model

To make that change, we can start with a model that has already learned to process language and adapt it to score decisions. Open-Jev does this with a pretrained decoder, so it gives us a direct comparison with the LLM we just followed. Laya starts with an encoder, which lets us see how a different backbone can serve the same task.

### Open-Jev: adapting a decoder

Open-Jev keeps **component 1**, the pretrained model’s tokenizer and embeddings, and **component 2**, its language-processing backbone. Together, these components interpret the request and relate it to the possible answers.

Open-Jev replaces **component 3**, the vocabulary head, with a decision head that scores the supplied candidates. It converts those scores into probabilities without entering the answer-generation loop.

![Replace the output head](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/decision.png)

[Watch this animation](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/decision.mp4)

To get a score for each of our 4 departments, Open-Jev constructs a separate input for each proposed answer. Each input contains the customer’s message, the question and 1 candidate answer. The backbone processes each input, and the decision head reads the final position’s representation to produce 1 score. The software combines the 4 scores into probabilities. <sup>[6](#ref-6)</sup>, <sup>[7](#ref-7)</sup>

![Score each candidate with a shared model](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/candidates.png)

The new head needs training to produce useful scores. Open-Jev keeps the original backbone weights frozen and trains small additions, called LoRA adapters, alongside the decision head. The adapters adjust how the backbone processes the input, while the head learns to score candidates. <sup>[8](#ref-8)</sup>

Replacing the head is 1 way to get these scores. Eric Zhang’s [openjev-sglang](https://github.com/ekzhang/openjev-sglang) makes a lighter adaptation: it keeps component 3 and reads the scores of single-token labels such as A, B, C and D. It avoids continuing the answer, but still computes the vocabulary output. A decision model can therefore retain the vocabulary head, though it still pays the cost of computing those scores.

### Laya: adapting an encoder

Both of those examples start with a decoder, but scoring supplied answers doesn’t require a model built for continuing text. Laya’s authors start with a pretrained encoder. Their model still uses **component 1**, tokenization and embeddings, and **component 2**, language-processing layers with attention. Here, component 2 uses context from both directions in the input, and a decision head fills the role of **component 3**.

[Laya’s English models](https://github.com/NandhaKishorM/laya) start from ModernBERT, and their input includes the text, the question and marked answer options.

After the encoder processes them, additional trained layers prepare the representations for scoring. A shared scoring network reads the positions marking the options and produces their scores. So, for our support example, all 4 department descriptions can be represented within 1 question’s input. <sup>[9](#ref-9)</sup>

Attention remains in this approach too, both in the pretrained encoder and in Laya’s added processing layers. Laya avoids generating an answer, while retaining the language-processing layers needed to interpret the text and score the options.

![Score marked options with an encoder](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/encoder.png)

| Component | Generative decoder LLM | Decoder decision model, as in Open-Jev | Encoder decision model, as in Laya |
| --- | --- | --- | --- |
| 1. Tokenization and embeddings | Convert text into model inputs | Reuse from pretrained decoder model | Reuse from pretrained encoder model |
| 2. Processing backbone | Decoder layers use current and preceding context | Retain decoder layers, adapted for decisions | Use encoder layers with context in both directions, plus trained processing layers |
| 3. Output head | Score vocabulary tokens | Score supplied candidate inputs | Score marked answer options |
| Answer-generation loop | Repeat for each output token | Removed | Absent |

### How this compares with BERT classifiers

If you’ve fine-tuned BERT for classification, a lot of this probably looks familiar. A BERT classifier could already read our support request and predict a department without generating a sentence. In fact, Laya builds on ModernBERT, so there’s a direct connection to the models many ML engineers have worked with.

Remember, though, that a conventional deep-learning classifier often learns a fixed set of labels for a particular task. A model trained to route support tickets has those department labels built into its output head. With these decision interfaces, the question and possible answers arrive with the request. We can ask which department should handle the message, then ask whether it contains a refund request, through the same interface. That makes this a broader way to use classification within an application than training a separate fixed-label head for each task.

There are precedents for that flexibility too. 0-shot natural-language inference models could evaluate candidate descriptions at runtime years before Jev. The [BART-MNLI example](https://huggingface.co/facebook/bart-large-mnli) turns a possible label into a statement and evaluates whether the input supports it. So I see these decision models as building on familiar classification ideas and making them easier to use across different application decisions. The ability to pass in questions and choices helps explain the appeal, but we also need to look at what the returned probabilities mean when software acts on them.

## 3. Training for calibrated decisions

The decision heads we’ve looked at produce scores that we turn into probabilities. Once software uses those probabilities to decide whether to act, the model’s estimate of certainty becomes part of the application’s behavior. It needs to be evaluated alongside whether the model chooses the right answer.

For our routing example, choosing billing is only part of the output. Assigning billing a 60% probability and assigning it 99% communicate very different levels of certainty. Across many comparable support-routing predictions assigned 80%, about 80% should be correct. That is what calibration means; it isn’t a guarantee about 1 individual request.

![What 80% means across predictions](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/calibration.png)

[Watch this animation](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/calibration.mp4)

A generative language model starts by learning to predict tokens, so its next-token probabilities describe what it is likely to write. Restricting the output to answer labels doesn’t establish that those probabilities match how often its decisions are correct. Later training may use human preferences to favor helpful responses, or verifiable rewards to favor answers that pass a check. Those approaches are commonly called RLHF and RLVR.

TypeSafe trains Jev for this probability-estimation task using what it calls **Reinforcement Learning for Calibrated Decisions**, or RLCD. Its stated training target is decisions accompanied by probabilities that reflect how likely they are to be correct. <sup>[10](#ref-10)</sup>

In reinforcement learning, the model is updated to favor outputs that earn better rewards. For calibrated decisions, the reward needs to account for the probability estimates, including being confidently wrong. TypeSafe has described that objective but hasn’t published the exact reward or update algorithm, so we can’t reconstruct Jev’s training from the name alone.

![Train probabilities against outcomes](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/training.png)

For an inspectable implementation, we can return to Laya. Its training materials describe probability-scoring rewards and policy-gradient updates, as well as fitting calibration settings using held-out examples. These are Laya’s implementation choices; we don’t know whether Jev uses the same methods. Laya reports overconfidence before calibration and improvement after task-specific training. <sup>[11](#ref-11)</sup>

Older classifiers can also be calibrated, and probability-scoring objectives can be used in supervised learning. To compare them with Jev, we’d need to test both on the same questions, check their probabilities against actual outcomes, and measure their serving costs.

## 4. Why decision models can be faster and cheaper

Once a model has learned to score the decisions we need, it can return those scores without writing an answer. The architecture changes we followed earlier can save computation in 2 ways:

- **Removing the answer-generation loop.** The model processes the input and scores the choices, then stops. It doesn’t keep running the decoder to write the answer token by token. This is the biggest saving when the alternative is generating a long answer.
- **Replacing the vocabulary head.** A language-model head projects the hidden representation into scores for the entire vocabulary, which can contain 100,000 or more tokens. A decision head only needs to score the candidates. We still convert those scores into probabilities, but over a much smaller set.

For the routing task, we need scores for 4 departments. If we ask an LLM for “billing” plus an explanation or a JSON response, generating that text adds computation. A decision model can return the scores for software to package. The free-form customer reply is a different task: we’d still use a generative model when the application needs to write it.

These changes can reduce the GPU time spent on each request and let the same hardware handle more requests. A smaller pretrained backbone can reduce the computation and memory requirements further.

A decoder-only LLM already has no encoder to remove. Calling a model “decoder-only” therefore doesn’t mean we’ve halved its computation. The savings above come from changing how it produces an answer.

The saving also depends on the workload:

- Long documents still require input processing through attention and the other layers.
- Large candidate sets can reduce the advantage, especially when an implementation repeats the input for every candidate.
- Avoiding a long generated response saves more than avoiding a single-token answer.

### Parallel sampling: answering several questions together

So far, we’ve asked the model to choose a department. The same customer message could also tell us about urgency and whether a refund was requested. A generative LLM returning all 3 in JSON normally writes the result as 1 sequence of tokens. Even though the questions are independent, the output is still generated sequentially.

TypeSafe describes Jev’s **parallel sampler** as producing the output probabilities for all questions in a single query. The department answer doesn’t have to be written out before the urgency answer can be returned. <sup>[1](#ref-1)</sup>

The open implementations illustrate ways to compute scores together: candidate sequences can be batched in the decoder approach, and question inputs can be batched in the encoder approach. Batching lets the hardware process multiple inputs concurrently. TypeSafe hasn’t published enough detail to equate its sampler with either implementation, or to say that all questions share 1 internal pass.

The model still uses attention to process the inputs. It can then return the decisions together, without writing each answer token by token.

![Return several decisions together](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/parallel.png)

[Watch this animation](https://aishreganti.com/blog/how-jev-works-under-the-hood/out/article/parallel.mp4)

## 5. Where I think decision models are heading

I do think Jev is a very smart way to package a capability. A lot of the interest makes sense when you realize how many application tasks can be expressed as decisions over a known set of choices. By changing the output space and training a model for those decisions, we can avoid generating an answer token by token and make many of these tasks faster and cheaper.

That starts to look useful in places where calling a generative model for every small decision would be too slow or expensive. [Made with Jev](https://madewithjev.com/) is a nice place to see the art of the possible: people are trying these models in routing workflows, browser agents and real-time applications. Looking at that range, I can see why builders are excited about having this kind of capability available through an API.

Over time, though, I do think the wave of standalone decision models might be absorbed by larger model providers. They have a reason to offer cheaper ways to handle the many requests that don’t need a generated response, and I wouldn’t be surprised to see them introduce their own versions in the near future.

As orchestration improves, a provider could offer decision models alongside its generative models and let an application choose what each step needs. Our support workflow could use a decision model to route a request and the other to write a reply, without sending every step through the same expensive model. That’s how I can see this becoming part of a larger model ecosystem, with cheaper options for routine decisions available through the platforms developers already use.

## References

1. <span id="ref-1"></span>[TypeSafe AI: Introducing System One Models and Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
2. <span id="ref-2"></span>[Vaswani et al.: Attention Is All You Need](https://arxiv.org/html/1706.03762v7)
3. <span id="ref-3"></span>[Devlin et al.: BERT](https://arxiv.org/abs/1810.04805)
4. <span id="ref-4"></span>[Qwen3.5-2B model card](https://huggingface.co/Qwen/Qwen3.5-2B)
5. <span id="ref-5"></span>[Hugging Face Transformers: How caching works](https://huggingface.co/docs/transformers/main/cache_explanation)
6. <span id="ref-6"></span>[Open-Jev: candidate input construction](https://github.com/Zefan-Cai/Open-Jev/blob/main/jev/api.py)
7. <span id="ref-7"></span>[Open-Jev: model implementation](https://github.com/Zefan-Cai/Open-Jev/blob/main/jev/model.py)
8. <span id="ref-8"></span>[Open-Jev: training and released models](https://github.com/Zefan-Cai/Open-Jev)
9. <span id="ref-9"></span>[Laya: architecture implementation](https://github.com/NandhaKishorM/laya/blob/main/laya/common.py)
10. <span id="ref-10"></span>[TypeSafe AI: machine learning primer and RLCD](https://docs.typesafe.ai/introduction/machine-learning-primer)
11. <span id="ref-11"></span>[Laya: training and calibration](https://github.com/NandhaKishorM/laya)
