This post is the short version of my video. Watch it here: I Tried to Fool 4 AI Decision Models With 1000 Trick Questions
A decision model does not write. You hand it a text and a question with a fixed set of answers, and it returns a probability for each answer. Routing, moderation, “is this tool call allowed”: that kind of job.
There are now several of them, they share one request format, and I could not find a comparison that did more than report one accuracy number. So I wrote my own test set, 1000 questions that each try to trick the model in a specific way, and ran four models on it: Jev (TypeSafe, hosted), Nimble 9B and tev1 4B (local, through Ollama) and Strands Decider 2B (local, its own server).
Everything is in the repository: github.com/huseyinbabal/can-you-fool-jev. Questions, labels, the runner, every request body that was sent and every answer that came back.
First, what this is not
It is not a benchmark. One person wrote the questions and the labels, each model ran once with default settings, and 843 of the 1000 questions are customer messages. I list 42 labels I consider debatable myself.
What makes it useful anyway is that the four models received byte-for-byte the same requests, except for the model field. So differences between them are differences between the models, not between prompts.
One request, three kinds of question
All four speak the same /v1/systemone format. A request carries the text to judge and any number of questions:
{
"state": "Oh great, the third 'replacement' charger also died after two days. Fantastic work, really. I need it working before my flight on Friday.",
"questions": {
"sentiment": {"type": "choice", "instructions": "What is the customer's real sentiment?",
"criteria": {"positive": "", "neutral": "", "negative": ""}},
"urgent": {"type": "noul", "instructions": "Is there a deadline?",
"criteria": {"true": "the customer needs it by a specific time", "false": "no time pressure"}},
"anger": {"type": "score", "instructions": "How angry is the customer, from 1 (calm) to 10 (furious)?",
"criteria": ["1","2","3","4","5","6","7","8","9","10"]}
}
}
A choice picks one option, a noul is a yes/no probability, a score is a number on a scale. The dataset has 463 choice questions, 321 yes/no and 216 scores.
The ranking
| Model | Accuracy | Confident and wrong | Median time |
|---|---|---|---|
| Jev | 91.4% | 7 | 275 ms |
| Nimble 9B | 81.8% | 36 | 400 ms |
| tev1 4B | 76.8% | 19 | 220 ms |
| Strands Decider 2B | 64.7% | 4 | 77 ms |
“Confident and wrong” counts wrong answers given with at least 90% probability. Jev’s time includes the network round trip; the other three ran on my Mac.
559 questions were answered correctly by all four. 40 fooled every one of them. If I stopped here you would learn almost nothing you could act on, so here is where the gaps actually are.
Finding 1: scores are a different sport
| Answer type | Questions | Jev | Nimble 9B | tev1 4B | Strands 2B |
|---|---|---|---|---|---|
| choice | 463 | 94% | 88% | 85% | 78% |
| yes / no | 321 | 94% | 81% | 83% | 74% |
| score 1–10 | 216 | 82% | 69% | 49% | 22% |

On choices the spread between best and worst is 16 points. On scores it is 60. A score counts as correct when it lands in a small accepted range (for example 7 to 9), and the two smallest models miss that range most of the time.
If your use case is “rate this from 1 to 10”, an overall accuracy figure tells you very little. Test scores separately.
Finding 2: each model has its own blind spot
Every question is tagged with the trap it sets. These are the traps where the models diverge most:
| Trap | Questions | Jev | Nimble 9B | tev1 4B | Strands 2B |
|---|---|---|---|---|---|
| Tone says one thing, content another | 20 | 90% | 50% | 30% | 10% |
| Severity hidden in a calm message | 28 | 79% | 57% | 43% | 39% |
| Sarcasm | 23 | 91% | 52% | 52% | 30% |
| Deadline buried in the text | 16 | 94% | 69% | 69% | 19% |
| Numbers you have to work out | 13 | 100% | 77% | 62% | 38% |
| Prompt injection inside the input | 26 | 92% | 77% | 81% | 62% |

“Tone versus content” is the clearest case: a polite, calm message reporting a total outage, or an all-caps message about a typo. The smaller models follow the tone.
These groups are small, 13 to 28 questions each, so treat the exact percentages as rough. The pattern is consistent, though: the gap between models is far wider on these than on the overall number.
Two of the 40 questions no model got right, with the expected label and what came back:
| Input | Expected | All four answered |
|---|---|---|
| “Can’t wait to spend my Saturday on the phone with you again!” | angry | happy |
| “Dear Sir, thank you for your prompt reply of 14 days.” | angry | neutral |
And one where I think the models are right and my label is wrong: “I’m really upset, my order never came, and it was a gift.” I labelled it sad. All four said angry.
Finding 3: is a low confidence worth anything?
A decision model returns probabilities, which is only useful if they mean something. Two questions: when a model is very sure, is it right? And when it is unsure, is that a real warning?
| Model | Answers at ≥ 90% | Right | Answers below 50% | Right |
|---|---|---|---|---|
| Jev | 630 | 99% | 116 | 75% |
| Nimble 9B | 587 | 94% | 184 | 63% |
| tev1 4B | 445 | 96% | 213 | 44% |
| Strands 2B | 103 | 96% | 266 | 27% |

High confidence is trustworthy for all four. The interesting column is the last one. When Strands is unsure it is right about one time in four, so its low confidence is an honest signal: you can route those cases to a person or a bigger model and lose little. It also almost never commits: only 103 of its 1000 answers reach 90%.
Nimble is the one to watch. It is confidently wrong 36 times, more than the other three together.
Finding 4: the fastest and the most accurate are different models
Strands answers in 77 ms on a laptop and gets 64.7%. Jev gets 91.4% and takes 275 ms including a trip over the network. Whether 200 ms matters depends entirely on where the decision sits: inside a typing indicator it does not, inside a loop that gates every tool call of an agent it might.
There is also one category where the hosted model is not the best. On the 24 product questions (voice commands, browser elements, SQL columns) Nimble and tev1 reached 96% and Jev 92%. With 24 questions that is one question of difference, but it is a reminder that “most accurate overall” is not “most accurate for you”.
Run it on your own data
The runner works with anything that exposes the same endpoint:
ollama pull nimble && ollama pull tev1:4b
python3 scripts/run_systemone.py nimble
python3 scripts/run_systemone.py tev1:4b
SYSTEMONE_HOST=https://api.typesafe.ai SYSTEMONE_KEY=$JEV_API_KEY \
python3 scripts/run_systemone.py jev-latest
python3 scripts/summary.py
A small set of your own data will tell you more than my 1000 questions. A few dozen real inputs from your system, including the awkward ones that already caused trouble, is enough to see which model fails where. The format of data/questions.csv is documented in the README, and runs are resumable.
If you disagree with a label, the repository takes issues.