
A System One model answers typed questions. You send it some text and a question with a fixed shape: pick one of these options, give a score from a scale, or answer true or false. It returns the answer with a probability for each option. It writes no free text.
We tested three of them as checkers for a document verifier, and added one chat model as a reference. This post reports two things: how often each one made the wrong decision on the same inputs, and how much input each one accepts.
The models
- JEV (Typesafe), called through OpenRouter's decisions endpoint.
- Kev 4B, called through the same endpoint.
- Strands Decider 2B, an open model from AWS Strands Labs, served on our own server.
- Amazon Nova 2 Lite, a general chat model, called through OpenRouter. It is a reference row. It answers the same question in free text with a quote and a one-sentence reason.
The three System One models take the same request: a state, a set of typed questions, and a model name.
What a decision is
A decision is one claim from a written answer about a contract, with the passages of that contract it should be checked against. The model chooses one of six verdicts:
- supported: the passages say what the claim says.
- unsupported: the passages say something different or nothing backs the claim.
- wrong clause: the claim takes its content from a clause that belongs to a different provision.
- missed: the claim says the contract has no such clause, but a passage is that clause.
- reversed: the claim gives a right or duty to the wrong party.
- not in passages: the passages do not cover the claim, so it cannot be checked.
We wrote the correct verdict for every claim before running anything. The set has 127 claims from labelled cases built on CUAD and MAUD, two public contract datasets. Rule-based checks settle some claims without a model. We counted only the claims that went to a model, and every model received identical claims and identical passages.
Each model gets one request per document, containing all of that document's claims and up to four passages of 2,500 characters for each claim. There was no fallback. If a model could not take a request, those claims stayed unanswered.
Wrong decisions
| Model | Correct | Wrong | No answer |
|---|---|---|---|
| JEV | 104 | 14 | 9 |
| Nova 2 Lite (chat) | 104 | 14 | 9 |
| Kev 4B | 91 | 32 | 4 |
| Strands Decider 2B | 6 | 69 | 52 |
A wrong decision is a verdict different from the correct one. We split the wrong ones into three types:
| Model | Failure, wrong label | Missed an error | False alarm |
|---|---|---|---|
| JEV | 7 | 4 | 3 |
| Nova 2 Lite | 6 | 5 | 3 |
| Kev 4B | 17 | 12 | 3 |
| Strands Decider 2B | 41 | 4 | 24 |
"Failure, wrong label" means the model saw a problem and named the wrong kind. "Missed an error" means it said supported where the claim was wrong. "False alarm" means it flagged a correct claim. "No answer" means the model said the passages do not cover the claim, or the request was rejected.
On the 33 claims held out from the tuning of our prompts, correct counts were 27 for JEV, 24 for Nova, 23 for Kev 4B, and 3 for Strands.
By kind of claim
Correct decisions, out of the number of claims of that kind:
| Claim should get | Claims | JEV | Kev 4B | Strands | Nova |
|---|---|---|---|---|---|
| supported | 76 | 66 | 73 | 0 | 67 |
| missed | 33 | 25 | 10 | 1 | 26 |
| unsupported (invented protection) | 7 | 7 | 7 | 0 | 7 |
| wrong clause | 5 | 5 | 0 | 5 | 1 |
| reversed | 4 | 0 | 0 | 0 | 1 |
| unsupported (misread) | 2 | 1 | 1 | 0 | 2 |
- Kev 4B has the highest count on supported claims, and it said "supported" 85 times against 76 that deserved it. It found 10 of 33 missed clauses, and it answered 0 of 5 wrong-clause claims correctly.
- JEV and Nova are close on every kind of claim.
- All four struggled with reversed claims. There are only four of them, so this is a signal to look at, not a rate.
Context window
We measured how much input each one accepts, by sending text of growing size until the request failed.
| Model | Accepted | Rejected |
|---|---|---|
| JEV | 25.8k tokens | about 34k tokens (error: max tokens exceeded) |
| Kev 4B | 7.7k tokens | about 8.5k tokens (error: parameter invalid) |
| Strands Decider 2B | 3.5k tokens | 5.2k tokens (the server reports a 4,096-token window) |
We found these limits by testing. Token counts for rejected sizes are estimated from the same text at smaller sizes, and we did not locate the exact cutoff for JEV or Kev.
In this evaluation JEV and Kev took every request. The Strands server rejected 26 of 102 requests with "prompt exceeds the context window of 4096 tokens". Those requests held 52 of the 127 claims, and none of them received a verdict.
A document check sends a whole contract's worth of passages, so window size sets how many claims a model can be asked about in one request. In production a request that is too long falls back to the chat judge, so a short window shifts work to the fallback.
Strands Decider
The 52 unanswered claims explain part of the Strands result. On the 75 claims it did judge, it answered "wrong clause" 70 times, "supported" 4 times and "missed" once. That is one label for nearly everything. Six of its 127 decisions were correct.
This run does not show how the model does on the task it was built for. The window is short, the six-way verdict question may be a poor fit for a 2B model, and we did not try a simpler question such as yes or no.
Limits of this test
- 127 claims. Several kinds have under ten claims, so read those columns as counts.
- We wrote the correct verdicts and the claims. The verdict descriptions given to the models came from a prompt tuned on part of this set, so the held-out figures are the cleaner comparison.
- The test measures a verdict against a fixed correct answer. It does not measure how the check performs inside a larger system.
- The Strands server was configured with a 4,096-token window. A different configuration could change its result.
- One run per model. We did not repeat runs.
Per-claim verdicts for all four models are in system-one-judge-decisions.csv.