This post is the short version of my video. Watch it here: LLMs talk. Jev decides.
Over the last few years we’ve thrown an LLM at almost every problem. Which team should get this support ticket? LLM. Is this comment harmful? LLM. What should this game character do next? LLM.
And the LLM politely writes a paragraph. “Certainly! Upon reviewing this ticket…”
None of these questions needs a paragraph. The answer is A, B or C. Yes or no. Attack or flee. It’s a decision.
The problem with asking an LLM
Send a ticket like “My Android timeline is blank!” to an LLM and ask which team owns it. The model generates its answer token by token, and every token waits for the one before it. Sometimes you get “Hmm, I think this might be a billing issue…” when all you wanted was one word: bug.
Structured output modes help, but the model still writes its JSON piece by piece. So the three problems stay:
- Slow: in TypeSafe’s comparisons, frontier LLMs take anywhere from 3 to 329 seconds on tasks like this.
- Expensive: output tokens usually cost several times more than input tokens.
- Unreliable: the output space is infinite, so anything can come back.
System 1 and System 2
The framing that made it click for me is Daniel Kahneman’s Thinking, Fast and Slow. System 1 is fast and automatic: two plus two, you just know. System 2 is slow and careful: seventeen times twenty-four, you have to stop and work it out.
LLMs behave like System 2. They reason, explain and write. But most jobs in automation (classification, routing, scoring, moderation) are System 1 jobs. We hired a philosopher for a job that needs a reflex.
What Jev is
TypeSafe calls Jev the first “System One model”. The best way to think about it: it’s not a chatbot, it’s a function call.
You give it unstructured state (a ticket, a chat history, a game’s JSON state) and typed questions. It returns typed values, each with a probability. In TypeSafe’s words: “Unstructured state in, typed probabilistic decisions out.”
It’s fast because it’s non-autoregressive. It doesn’t generate tokens at all; it decides every field in one pass, in parallel. TypeSafe quotes 70 to 500 ms end to end, and asking several questions costs barely more than asking one.
There are three question types:
- Choice: “Which team owns this ticket?” You get a probability for every option, up to 255 options.
- Score: “How urgent is it, from 0 to 10?” You get a distribution over every level, not one number.
- Yes or no: “Is this request harmful?” You get a probability between 0 and 1.
No type errors, but not no wrong answers
The boldest claim is that Jev can’t make a type error. An LLM asked for bug can return Bug., BUG, looks like a bug, half a JSON object or null. Jev’s output space is only the type you defined. If the options are billing, bug and account, there is no fourth option.
The price is that Jev doesn’t generate free text. And “no type errors” doesn’t mean “no wrong answers”: it can still pick the wrong option. That’s why the probabilities matter.
In code
Pydantic AI has a ready-made integration:
pip install "pydantic-ai-slim[typesafe]"
from typing import Annotated, Literal
from pydantic import BaseModel
from pydantic_ai import Agent, BoolCriteria
class Ticket(BaseModel):
area: Literal['billing', 'bug', 'account']
urgent: Annotated[bool, BoolCriteria(
true='The customer is losing money or has a deadline today.',
false='It can wait its turn in the queue.')]
app: Literal['web', 'ios', 'android'] | None
agent = Agent('typesafe:jev-latest', output_type=Ticket)
r = agent.run_sync("Android timeline blank since the update, standup in 10 min!")
No prompt engineering, no “please return only JSON”. You get a plain Python object back, and the response also carries a confidence for every field. Spelling out what “yes” and “no” mean with BoolCriteria matters, because Jev reads text very literally.
Confidence is your seatbelt
This is the part I find most useful in practice. Ask Jev first. If the confidence is above your threshold, act automatically, in milliseconds and almost for free. If it isn’t, hand that case to something slower and more careful: a big LLM, or a human. Pydantic AI has FallbackModel and decision thresholds for exactly this.
Jev takes the easy majority, and you only pay for the expensive model when you need it.
The numbers, and a caveat
TypeSafe lists 4.2 cents per million input tokens, and output is free because there is no generated text. They say it’s up to 200 times faster and up to 400 times cheaper than frontier LLMs.
Those numbers come from the company’s own tests. Test it on your own workload.
Where I wouldn’t use it
The docs are honest about this:
- It doesn’t generate text. A string field falls back to an LLM.
- Hard limits: 255 options per question, a 32k token context.
- It’s unreliable at arithmetic and counting. Leave that to code.
- It struggles with indirect, multi-hop questions.
- It’s vulnerable to prompt injection hidden in the text.
- No images yet.
The rule I took away: Jev decides, code computes, LLMs think.
Why “Jev”?
The name comes from William Stanley Jevons. In 1865 he noticed that as steam engines used coal more efficiently, coal consumption didn’t go down. It exploded. That’s TypeSafe’s bet: once machine intelligence gets cheap enough, it ends up everywhere.
LLMs and Jev aren’t rivals. The best systems will use both.