How Jev Helped Us Cut Voice-Agent Decision Time by 65%
Reading Time: 8 minutes
Ali Kemal Coşkun
AI Software Engineer
Eight points more accurate than the strongest language model a voice agent can afford to run, and a decision in a third of the time. 616 decisions, Turkish and English, measured inside a multi-agent voice platform.
At a Glance
- A call in Phoebe, Commencis’s voice AI platform, passes through several agents before it reaches an answer.
- Every turn begins with a decision rather than a sentence: who takes the caller, what step follows, which tool to call, or whether the right move is simply to ask.
- That decision has always belonged to the language model. In this evaluation it belonged to Jev.
- Jev chose the right move 95.1% of the time, against 87.3% for the strongest model a voice agent can afford to run. It decided in 331 ms, 2.9× faster than the eight language models’ 955 ms median.
- The language model keeps the language. It writes every word the caller hears, and the occasional free-text field a chosen tool needs, which takes Jev from 95.1% to 95.6%.
One turn, two jobs
Every turn in a call contains two different problems. The first is a decision: given everything that has happened, what should the assistant do now? The second is language: if the assistant is going to speak, what exactly should it say?
The conventional design assigns both to the same language model. The cost of doing so is easy to overlook, because the decision never appears in the output: the caller hears only the sentence.
In Phoebe the decision is a choice among a small set of moves: speak to the caller, hand the call to another agent, give the line back to the agent that delegated the task, advance to the next step, or call a tool. The set is never open-ended. At any moment, the agent’s configuration fixes which tools exist and which agents can be reached.
That structure is what made a decision model worth evaluating. TypeSafe AI introduced Jev in September 2026 as the first of what it calls System One models, a name the company takes from the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning, and credits to Kahneman’s Thinking, Fast and Slow. Instead of producing text token by token, the model evaluates typed questions against a state and returns typed values with probability distributions, drawn from an answer space defined in advance. Here that answer space is the set of moves available at that instant.

Figure 1. Phoebe supplies the decision context, applies whatever comes back, and asks the language model for the words. A logical view, not a deployment diagram.
A handover can be a detour, not an ending
Calls that cross several agents make these decisions considerably harder.
Consider a lost card. A greeting agent routes the call to banking. Banking brings in the lost-card specialist. The specialist cannot block anything until the caller’s identity is confirmed, so it hands the line to an identity-check agent. When the check succeeds, control has to come back to the specialist, and the original task has to continue where it stopped.
Every arrow in that sequence is a decision, and so is every tool call and every question along the way. Each one can fail in a way the caller notices: a transfer before they have explained the problem, a tool call before a necessary question, or a subtask that never gives the line back.
Scoring the decision, not the whole call
To see the decision clearly we measured it on its own, rather than scoring whole calls and inferring what happened inside them.
We started from a real Phoebe agent setup covering banking, telecom, insurance and travel. From it we generated 52 sample calls and split them into 308 decision points. A decision point is the snapshot of the instant before an agent acts: the conversation so far, the call state and the step’s instruction, and the tools and agents available right then. Each point carries one expected move, used only for scoring and never shown to a model.
Every point exists in Turkish and in English, which gives 616 scored evaluations. They are 616 decisions, not 616 separate calls, and none of them are recordings of real customer conversations.
We compared three configurations:
- Jev. The decision model chooses the move on its own.
- Jev + fill. Jev chooses the move; when the chosen tool needs a free-text field, such as the subject of a complaint, a small language model writes that field and nothing else.
- Language models. A general-purpose model chooses the move and its arguments, which is how Phoebe worked before. We evaluated eight models: GPT-6 Sol and Luna, GPT-5.4 mini and nano, GPT-4.1 mini, Claude Haiku 4.5, Gemini 2.5 Flash and Flash Lite, with thinking switched off wherever the provider allows it.
What we found

Figure 2. Correct move, and the stricter score that also requires every argument to be right.
In absolute terms, Jev made 30 incorrect moves across the 616 evaluations. Claude Haiku 4.5, the strongest of the models a voice agent can afford to run, made 78.
One language model came close on accuracy. GPT-6 Sol reached 93.3%, within two points of the decision model, at a median of 1,545 ms per decision. A voice turn has to hold speech recognition, the decision, the reply and speech synthesis inside roughly a second and a half before the pause becomes audible, and the working guidance for the model stage is around 600 ms. A decision of that length consumes the turn before a word is synthesised, which is why Sol belongs in the table as an upper bound rather than in the call path.

Figure 3. Median time per decision, and which stage of the turn it covers.
The fill model ran in 30 of the 616 decisions, about one in twenty, and added a median of 719 ms on those. The 336 ms median describes the whole set, not the decisions that needed fill.
Put the two measures on one plot and the shape of the trade-off is visible at once. The general-purpose models sit in a band around 85% at roughly a second per decision; the decision model sits ten points higher at a third of that time; the small open-weight models are faster still and far less accurate.

Figure 4. Accuracy against median decision time.
Where the gap comes from

Figure 5. Accuracy by the kind of move the moment called for.
The widest gap concerns knowing when not to act. When the right move was to speak, usually to ask the question the step depends on, Jev chose it 98.9% of the time and GPT-4.1 mini 73.9%. That single category accounts for 47 of its 87 wrong moves. In the error traces the model calls a tool before asking, or connects an agent before the caller has provided enough information to route on.
Handovers follow the same pattern in both directions: choosing the right agent, and giving the line back when a subtask is done. Returning the line is the weakest category for every configuration, and it is the one a caller notices most directly, because the alternative is silence or a repeated question.
The pattern is not universal, and the exception should be stated plainly. The language model is slightly better at picking the exact tool, and tool choice accounts for 14 of Jev’s own 30 errors. On step transitions the configurations are level.
A further finding accompanies the scores, and it is easier to read backwards. The bot runs live on full prompts that carry explicit handover and routing rules. Cutting those prompts to a third of their length, which is the compact set these results were obtained on, costs GPT-4.1 mini 5.7 points and Jev 0.7. The decision context can therefore stay lean when Jev is the one deciding, and the budget that is freed can go to the conversation instead. This is not a claim that detailed instructions are unnecessary, and it says nothing about what those rules do for the words the caller hears.

Figure 6. The same decisions under the full prompts and under prompts a third of their length.
Keeping the language model where language is required
The result is a hybrid architecture, not a voice agent without a language model. In Phoebe the decision arrives as a typed move and then takes one of four paths: the language model writes the reply when the move is to speak, with tool calls disabled; Phoebe executes the move itself when it is confident and its arguments are known; the language model writes one free-text field when a chosen tool needs it, and Phoebe then runs the tool; and everything else, including every decision the model is unsure of, falls back to the language model as before.
The third path is the one the scores justify. On the decisions where an argument was checked, Jev alone was right 72.8% of the time, Jev with fill 90.6%, and GPT-4.1 mini 98.3%. A model that returns choices cannot write a complaint summary, and the remedy is to let it choose and let the language model write.

Figure 7. Argument accuracy on the decisions where an argument was checked.
One implementation detail deserves mention, as a live call exposed it. Certain tools wait on the caller, such as a keypad entry or a confirmation. Executing such a tool directly, with no sentence attached, leaves the caller in silence while the tool waits. Each tool is therefore offered twice, once on its own and once as a move made while speaking, so that the assistant can request the customer number in the same turn that opens the keypad.
Open-weight alternatives
We also ran the same decisions through two hosted open-weight choice models. SemIf, a 4-bit quantized Qwen3.5-4B, selected the correct move in 54.5% of cases at a 231 ms median. Laya reached 31.2% at 61 ms. Both are fast and neither is sufficiently accurate for this workload, and the reason is more instructive than the scores.
A real decision point carries a full instruction, the conversation so far and a dozen tool descriptions. Laya cannot hold that in one question, so we compressed the input to fit. Compression alone does not explain the gap: given the same compressed inputs, Jev scored 95.1% and the language models between 75% and 93%.
The limitation is structural as much as numerical. Laya packs its options into a fixed budget that counts the option identifiers as well as their text, and Turkish labels consume more of that budget than English ones, so fourteen options fit in Turkish where twenty fit in English. Within that budget, fewer options carrying their descriptions did better than more options carrying bare labels, and translating the whole task into English changed nothing.
Neither model was fine-tuned for this task. Both model cards describe the released weights as a base to specialise, so training one of them on decision data of this kind remains the obvious next step, and an untested one.
Implications for voice agents
Voice agents have largely been built around a single model that listens, decides and speaks in one pass, paying for that pass on every turn. Separating the decision changes what each component is required to do well: a decision layer that moves the call through the correct agents and steps within a few hundred milliseconds, and a language model reserved for what only it can do, which is to hold a conversation with a person.
The value is not limited to a faster call. Routing, confirmations and handovers become a responsibility that can be measured and improved independently of the words. The remaining question is how much of the improvement survives the full voice loop, and answering it requires an end-to-end evaluation. This work establishes the means to examine that question one component at a time.
About Phoebe. Phoebe is Commencis’s voice AI platform. For how we built its Turkish text-to-speech engine, read Phoebe’s Voice: Building a Text-to-Speech Engine That Actually Sounds Turkish.
Reading Time: 8 minutes
Don’t miss out the latestCommencis Thoughts and News.
Ali Kemal Coşkun
AI Software Engineer
Commencis Thoughts
Commencis Thoughts explores industry trends, emerging technologies and global consumer culture through the eyes of Commencis leaders, strategists, designers and engineers.

