Skip to content
Huseyin Babal
Go back

LLMs talk. Jev decides.

This post is the short version of my video. Watch it here: LLMs talk. Jev decides.

Over the last few years we’ve thrown an LLM at almost every problem. Which team should get this support ticket? LLM. Is this comment harmful? LLM. What should this game character do next? LLM.

And the LLM politely writes a paragraph. “Certainly! Upon reviewing this ticket…”

None of these questions needs a paragraph. The answer is A, B or C. Yes or no. Attack or flee. It’s a decision.

The problem with asking an LLM

Send a ticket like “My Android timeline is blank!” to an LLM and ask which team owns it. The model generates its answer token by token, and every token waits for the one before it. Sometimes you get “Hmm, I think this might be a billing issue…” when all you wanted was one word: bug.

Structured output modes help, but the model still writes its JSON piece by piece. So the three problems stay:

System 1 and System 2

The framing that made it click for me is Daniel Kahneman’s Thinking, Fast and Slow. System 1 is fast and automatic: two plus two, you just know. System 2 is slow and careful: seventeen times twenty-four, you have to stop and work it out.

LLMs behave like System 2. They reason, explain and write. But most jobs in automation (classification, routing, scoring, moderation) are System 1 jobs. We hired a philosopher for a job that needs a reflex.

What Jev is

TypeSafe calls Jev the first “System One model”. The best way to think about it: it’s not a chatbot, it’s a function call.

You give it unstructured state (a ticket, a chat history, a game’s JSON state) and typed questions. It returns typed values, each with a probability. In TypeSafe’s words: “Unstructured state in, typed probabilistic decisions out.”

It’s fast because it’s non-autoregressive. It doesn’t generate tokens at all; it decides every field in one pass, in parallel. TypeSafe quotes 70 to 500 ms end to end, and asking several questions costs barely more than asking one.

There are three question types:

No type errors, but not no wrong answers

The boldest claim is that Jev can’t make a type error. An LLM asked for bug can return Bug., BUG, looks like a bug, half a JSON object or null. Jev’s output space is only the type you defined. If the options are billing, bug and account, there is no fourth option.

The price is that Jev doesn’t generate free text. And “no type errors” doesn’t mean “no wrong answers”: it can still pick the wrong option. That’s why the probabilities matter.

In code

Pydantic AI has a ready-made integration:

pip install "pydantic-ai-slim[typesafe]"
from typing import Annotated, Literal
from pydantic import BaseModel
from pydantic_ai import Agent, BoolCriteria

class Ticket(BaseModel):
    area: Literal['billing', 'bug', 'account']
    urgent: Annotated[bool, BoolCriteria(
        true='The customer is losing money or has a deadline today.',
        false='It can wait its turn in the queue.')]
    app: Literal['web', 'ios', 'android'] | None

agent = Agent('typesafe:jev-latest', output_type=Ticket)
r = agent.run_sync("Android timeline blank since the update, standup in 10 min!")

No prompt engineering, no “please return only JSON”. You get a plain Python object back, and the response also carries a confidence for every field. Spelling out what “yes” and “no” mean with BoolCriteria matters, because Jev reads text very literally.

Confidence is your seatbelt

This is the part I find most useful in practice. Ask Jev first. If the confidence is above your threshold, act automatically, in milliseconds and almost for free. If it isn’t, hand that case to something slower and more careful: a big LLM, or a human. Pydantic AI has FallbackModel and decision thresholds for exactly this.

Jev takes the easy majority, and you only pay for the expensive model when you need it.

The numbers, and a caveat

TypeSafe lists 4.2 cents per million input tokens, and output is free because there is no generated text. They say it’s up to 200 times faster and up to 400 times cheaper than frontier LLMs.

Those numbers come from the company’s own tests. Test it on your own workload.

Where I wouldn’t use it

The docs are honest about this:

The rule I took away: Jev decides, code computes, LLMs think.

Why “Jev”?

The name comes from William Stanley Jevons. In 1865 he noticed that as steam engines used coal more efficiently, coal consumption didn’t go down. It exploded. That’s TypeSafe’s bet: once machine intelligence gets cheap enough, it ends up everywhere.

LLMs and Jev aren’t rivals. The best systems will use both.

Sources


Share this post:

Next Post
What really happens when you kubectl apply