System-One Control Bench
How well do individual decisions add up to a completed task?
Some AI models write an answer. Others choose from a list of possible answers. This project tests whether such decision models can guide an agent through a small grid puzzle, and how they compare with chat models and simple programmed strategies.
The task is easy to describe: reach the goal, avoid walls, and collect a key when a locked door blocks the way. At each move, the model sees the puzzle as text and chooses its next move from a list. It then sees the new situation and chooses again, so to finish it has to keep making useful decisions over many moves. Because the puzzles are small, a solver knows the shortest route from every position. That lets us check both whether a model finishes a puzzle and whether each move was one of the best available; often several moves are equally good.
Does explaining the situation help? Every model receives the complete map, the rules, its position and whether it carries the key. We also test what happens when code explains parts of the situation for it: what is nearby, which moves it has already made, what each possible move would do, and which object to aim for next. These additions help separate two challenges: reading the map and choosing what to do.
What did we find? More detailed descriptions often help, but strong single decisions do not guarantee a finished game. With full context, Jev chooses an optimal move in about 90% of the positions of a separate exam, where every model answers the same fixed situations, yet it finishes only 57 of the 100 puzzles when it plays whole games. One test measures decisions in shared situations; the other measures whether a model can carry a task through to the end.
The puzzles, recorded games, code and analysis are public. This is a small, fixed benchmark for studying sequential decisions; the report explains the methods, results and limits.
Read the report · Explore the code and data · Submit a model
Watch two recorded games
Follow two games one move at a time to see how the players behave: where they make progress, take a detour or repeat an earlier mistake. An optimal move starts a shortest route from where the agent is, even if earlier moves already took it off the shortest route from the start. These are saved benchmark games, so playing them back makes no new model requests.
A agent · G goal · K key · D locked door · # wall
Jev, full context
DeepSeek V4.1 Flash (reasoning), full context
Leaderboard
Each player attempts the same 100 puzzles, one step per move, under two conditions:
- Map only: the rules, the complete map, the agent's position and whether it carries the key.
- Full context: everything in map only, plus the surroundings, the move history, the outcome of each move on offer and the next target.
Players are sorted by full-context games won. Games won counts the puzzles finished within the move limit, which is twice the shortest route. Mean progress measures how much closer the agent got to the goal at its best point in each game, averaged over the 100 puzzles: 1 means finished, 0 means it never got closer than where it started. Reported API cost covers both conditions together, 200 games; a dash means no cost was reported, and local computing is not included. The baselines are reference points: random moves, two greedy strategies and a solver that always chooses an optimal move.
| Player | Full context | Map only | Reported API cost (USD) | ||
|---|---|---|---|---|---|
| Games won | Mean progress | Games won | Mean progress | ||
| DeepSeek V4.1 Flash (reasoning) chat model detailsdeepseek/deepseek-v4.1-flash via OpenRouter (DeepInfra), reasoning capped at 1,024 tokens | 80 / 100 [72–88] | 0.86 [0.80–0.92] | 67 / 100 [58–76] | 0.79 [0.72–0.85] | 0.72 |
| Gemma 4 26B chat model detailsgoogle/gemma-4-26b-a4b-it via OpenRouter (DeepInfra), reasoning off | 61 / 100 [52–70] | 0.71 [0.64–0.79] | 35 / 100 [26–45] | 0.45 [0.37–0.54] | 0.12 |
| DeepSeek V4.1 Flash chat model detailsdeepseek/deepseek-v4.1-flash via OpenRouter (DeepInfra), reasoning off | 60 / 100 [51–70] | 0.72 [0.64–0.79] | 39 / 100 [30–49] | 0.55 [0.47–0.63] | 0.10 |
| Jev decision model detailsjev-1.13.0 through its API; cost not reported | 57 / 100 [47–67] | 0.68 [0.60–0.76] | 32 / 100 [23–41] | 0.43 [0.35–0.51] | — |
| Qwen3.5-4B chat model detailsQwen/Qwen3.5-4B in bfloat16, run locally, thinking off, scored on option-number probabilities | 45 / 100 [35–55] | 0.66 [0.59–0.73] | 14 / 100 [8–22] | 0.34 [0.27–0.41] | — |
| GLiClass decision model detailsknowledgator/gliclass-modern-large-v3.0, run locally on a CPU | 13 / 100 [7–20] | 0.22 [0.15–0.29] | 3 / 100 [0–7] | 0.10 [0.07–0.15] | — |
| Laya decision model detailsconvaiinnovations/laya, typed-decisions, run locally on a CPU | 12 / 100 [6–19] | 0.26 [0.21–0.33] | 2 / 100 [0–5] | 0.10 [0.07–0.15] | — |
| Programmed baselines (they do not read the request) | |||||
| Solver baseline detailsAlways plays an optimal move; the upper bound | 100 / 100 [100–100] | 1.00 [1.00–1.00] | 100 / 100 [100–100] | 1.00 [1.00–1.00] | — |
| Greedy (walls) baseline detailsStraight at the next target, never into a wall | 37 / 100 [28–48] | 0.47 [0.39–0.56] | 37 / 100 [28–48] | 0.47 [0.39–0.56] | — |
| Greedy baseline detailsStraight at the next target, even into a wall | 28 / 100 [19–37] | 0.36 [0.28–0.44] | 28 / 100 [19–37] | 0.36 [0.28–0.44] | — |
| Random baseline detailsA random offered move | 4 / 100 [1–8] | 0.25 [0.20–0.30] | 4 / 100 [1–8] | 0.25 [0.20–0.30] | — |
The full table adds SPL and the time per answer.
What information does a player receive?
The complete map is always available. Four optional components explain the situation in different ways:
| Component | What it tells the player |
|---|---|
| Surroundings | What is next to the agent, and how far away the key, door and goal are. |
| Move history | Every move the player has already made, and what it did. |
| Move outcomes | What would happen if the player chose each option. |
| Subgoal | Which object to aim for next: the key, the door or the goal. |
These descriptions are produced by code. They are real help, so the results measure the model together with the information it receives. The study tests ten conditions: the map alone, each component on its own, all four together (full context), and all four with one removed. This shows both whether a component helps by itself and whether it still matters when the others are present.
All ten conditions
| Condition | Name in the code | Added to the map |
|---|---|---|
| map only | map | nothing |
| map + surroundings | map+surroundings | surroundings |
| map + move history | map+memory | move history |
| map + move outcomes | map+lookahead | move outcomes |
| map + subgoal | map+subgoal | subgoal |
| full context | everything | all four components |
| full context minus surroundings | everything-surroundings | all but surroundings |
| full context minus move history | everything-memory | all but move history |
| full context minus move outcomes | everything-lookahead | all but move outcomes |
| full context minus subgoal | everything-subgoal | all but the subgoal |
Choosing several steps at once
The leaderboard uses one step per move, with four directions on offer. The wider study also tests moves of two or three steps chosen together, including rules where shorter moves stay available. Under these rules the player commits to a whole sequence before seeing the new situation: a blocked step is wasted, but the remaining steps still run, and reaching the goal ends the move. Longer moves change both the number of options and how far the agent goes before the player sees the board again.
| Rules | One move is | Options |
|---|---|---|
compass | one step north, south, east or west (the leaderboard's rules) | 4 |
two-moves | exactly two steps | 16 |
up-to-two-moves | one or two steps | 20 |
three-moves | exactly three steps | 64 |
up-to-three-moves | one, two or three steps | 84 |