In my last post Claude and I failed to get reinforcement learning to play Manic Miner, and we ended up solving the whole game with plain old search.
Around the same time, TypeSafe released Jev. Jev is what they call a âSystem Oneâ model - the name comes from the fast, instinctive kind of thinking, as opposed to slow, deliberate reasoning. It doesnât generate text, it makes decisions.
You give it some information about the situation (as JSON), a paragraph of instructions, and a list of options, and it tells you how likely it thinks each option is to be the right one. Itâs fast (around 270ms per decision for me) and cheap - a complete attempt at a cavern costs about half a cent.
Jev has already been used to play plenty of games. TypeSafeâs launch demos had it playing Doom from structured game state, and playing Wikiracing (get from one Wikipedia page to another using only links). People have since built Minecraft, Subway Surfers and driving demos, and thereâs a Minecraft speedrun agent that pairs a planner with Jev.
So: can Jev play Manic Miner? For anyone who missed the last post - you control Miner Willy, and in each of the 20 caverns he has to collect all the keys and then get to the exit, without touching an enemy, falling too far, or running out of air.
This has been a pretty interesting experiment. There is also a big caveat in this work - there are probably better ways to present the information to Jev and get better decisions. More research is required.
Itâs also entirely possible that we just arenât giving Jev good facts to work on. How do you describe a cave from Manic Miner in text?
At first I thought Jev was smashing it - performing much better than I expected.
Then I tested things using a very simple algorithm-based solution - this made me think that maybe Jev wasnât doing anything clever at all.
But, after a bit more investigation, I came back to the conclusion that Jev was actually doing something useful.
And then we fixed the information we were giving it, and the simple algorithm pulled ahead again.

Describing the game to Jev
Jev isnât multimodal - I canât just hand it a screenshot and say âwhat should Willy do?â. It also doesnât do reasoning or planning. It makes one quick judgement from whatâs in front of it.
My first thought was to give it a map of the cavern as text. This is an ASCII map of Central Cavern at the start of a run:
#........A.X....X............B.#
#...............C..............#
#..............................#
#..............................#
#......................XD..X...#
#=============~~~~=~~~~========#
#.............................E#
#===.........GG................#
#............GG..###.X.........#
#====...<<<<<<<<<<<<<<<<<<<<...#
#............................==#
#..............................#
#...........X.......###~~~~~===#
#.WW.===============.........PP#
#.WW.........................PP#
#==============================#
Each character is one cell of the screen. WW is Willy (heâs two cells wide and two high, and so are the guardians and the portal), A to E are the keys, P is the exit portal, G is a guardian, X is a nasty (the poisonous plants and stalactites), = is floor, ~ is crumbling floor, < is the conveyor, # is wall, and . is empty space. Jev gets a legend that explains all of this.
To a person, thatâs a pretty good description - you can trace the whole route on it. Jev canât.
You can see this when we ask Jev which key to go for next. It gets this map, plus a few facts about each key, like which side itâs on and how far away it is. Key D sits on the same top floor as keys A, C and B, and the only way up to that floor is at its far left end. So once Willy has collected E, a good order is A, C, D, B, walking right along the top floor.

Jev picks D, the key that looks nearest - on the facts, D is 5 cells away and A is 20. And it picks D every single time - in all 977 of our recorded runs that got this far, through nearly two weeks of changes to the harness. D gets a probability of about 0.8, and A about 0.04.

Going for D isnât a disaster - Willy has to go all the way round to the left either way. But Jev didnât see that. In one run, Willy collected E at decision 14, and it took until decision 67 before he found the way up at the left and got his next key (he picked up A, then C, on the way to D). When asked again at decisions 40 and 65, Jev was even more sure about D (0.92 and 0.90).
Claude (a reasoning model) worked it out from the same map: Willy can only reach the top floor from the far left, so he meets the keys in the order A, C, D, B as he walks right. Following a route across four floors takes several steps of reasoning, and Jev makes one quick judgement.
Itâs the same for moves. We tried showing Jev the map as it would look after each possible move, with no hint about which way to go. It finished 1 run of Central Cavern out of 20, and most runs got one key or none.
Reading a route off a map is not Jevâs strong point. We need to describe what is happening - including which way to go.
This description is what our harness is responsible for - all the code that sits between the game and Jev. Claude built it around the emulator and it does quite a bit of work. A slightly worrying amount of work to me.
- It reads Willy, the guardians, the keys, and the tiles out of the Spectrumâs memory.
- It tries all six moves (walk left/right, jump left/right/up, and wait) in the emulator. So we know exactly what each move does.
- It throws away any move that kills Willy or does nothing (i.e. Willy stays in the same place). It also throws away moves into a dead end: for each move, it plays up to four more moves ahead in the emulator to check that Willy can still stay alive - this is where it starts to feel like we are getting into the realm of cheating.
- It turns everything into facts relative to Willy, including some memory: âWilly has been here beforeâ, âWilly already tried this move from hereâ.
- For each move, it adds a fact called
progress: does this move take Willy ânearerâ to where he needs to go, or âfartherâ away?
To play the game, we run the cavern, collect all the facts and possible moves and then ask Jev to pick the best move.
{
"target": {
"what": "selected key", "side": "right", "horizontal_cells": 16,
"height": "higher", "way_up": { "side": "right", "horizontal_cells": 6 }
},
"to_the_left": { "first_thing": "nasty", "distance_cells": 1 },
"guardians": [ { "side": "same column", "height": "higher", "direction": "at Willy" } ],
"progress_measures": "distance to the way up",
"moves": {
"walk_right": { "movement": "Willy moves 1 cell to the right",
"progress": "nearer", "place": "new place", "tried_from_here": "no" },
"jump_left": { "movement": "Willy moves 4 cells to the left",
"progress": "farther", "place": "visited before", "tried_from_here": "no" },
"wait": { "movement": "Willy stays in the same place",
"progress": "same", "place": "visited before", "tried_from_here": "no" }
},
"moves_not_offered": {
"walk_left": "kills Willy: nasty",
"jump_right": "kills Willy: guardian",
"jump_up": "no effect: Willy stays in the same place"
}
}
Along with this comes a paragraph of instructions that explains the goal and what each fact means. Jevâs answer: walk_right 0.71, wait 0.18, jump_left 0.11. It picks walk_right.
Every so often, Jev also decides which key to go for next, as in the Central Cavern example above. progress is then measured towards whichever key it picked.
First impression: Jev is smashing it - this is amazing!
With all that in place, Jev completed Central Cavern in 10 runs out of 10. The Menagerie and a couple of the one-key caverns went well too. It felt pretty good - although, looking back, those are the easier caverns, and most of the caverns were never completed at all.
The map did help Jevâs key choices in a direct test. We asked it for the next key in 13 situations from the first four caverns. With the map and the facts together, it picked the same key as our completed runs in 12 of 13, against 5 of 13 with the facts alone. In Central Cavern, the map made it start with E instead of D. But better key choices didnât lead to more completed runs. When Claude chose the key order for the first four caverns and Jev did the moves, Jev didnât complete any more runs than with its own key choices.
Second thoughts: is Jev actually doing anything?
Working with Claude driving Jev was an interesting experience. Claude was able to analyse each run and determine where and why Jev got stuck. And Claude was able to suggest better facts to get it unstuck.
The problem is, every âbetterâ fact moves more of the âintelligenceâ into the harness.
The biggest one of these is progress. When the key is on the same floor as Willy, ânearerâ means nearer to the key. When the key is on a different floor, the harness works out a way up (or down) from the map, and ânearerâ means nearer to that spot.
After a while, progress starts to resemble a very obvious hint of the route to take through the cavern. Thereâs a real danger of solving the cavern using search and then presenting Jev with a bunch of moves along with a massive hint on which one to pick.
Instructions did seem to make a big difference. One of the difficulties here is that Claude does talk like an absolute lunatic sometimes, and it channels this into the instructions:
âIf Willy comes back to the same places again and again, the direct way is closed, and he must go a different way, even if that way goes away from the target first.â
Amazingly, this did actually seem to help!
So, weâve got a real danger of actually just being able to run the harness and then pick the âbestâ move presented by the harness without needing to use Jev at all.
This basic algorithm works very well:
- If a move completes the cavern or collects a key, pick it.
- Otherwise pick a move thatâs ânearerâ according to the harness.
- Otherwise pick any move.
This simple set of rules canât choose the key order, so we gave it a known-good key order for each cavern, and gave Jev the same order for comparison. We ran each cavern 10 times, and measured how far each run got: the share of the cavernâs keys and portal it reached. A complete run is 100%, and a run that collects 3 of 5 keys is 50% (3 of 6).
Across the whole game, Jev reached 31% of all the keys and portals, and the rule 26%. Both finished Central Cavern every time and The Cold Room most of the time. Jev picked a move the rule could also have picked in 93% of its decisions, against 54% for a random move. Take progress away and Jev almost never finishes a cavern. Random moves barely get started (4%).

So my conclusion at this point was: Jev can play Manic Miner, but the harness is actually doing all the work.
Why do so many caverns stop partway? Most of those runs didnât die - they got stuck, walking back and forth until 30 moves passed with nothing new. In several caverns Jev got about halfway and then stalled. The harness only knows the way to the next floor, and in caverns where the route winds across several floors, ânearerâ leads straight into a loop. Itâs highly likely that the âfactsâ we are presenting Jev work well on the first three caverns, because they are the ones we tested and tuned against.
Digging deeper
After some deep thinking from Claude, we tried some more experiments.
We added a second simple rule, just for keys - âgo for the nearest keyâ - and gave it to both Jev and the algorithm.
With the same key choice, Jev now got clearly further than the rule: 39% of all the keys and portals against 25%, and it finished caverns about twice as often. The same happened when both used the nearest-key rule (32% against 19%). Thatâs very unlikely to be luck - roughly a 1 in 1,000 chance.
On some caverns Jev is better
Jev stands out in two caverns, Wacky Amoebatrons and Amoebatronsâ Revenge: it finishes most runs there, and the algorithm almost never does. It also gets much further than the algorithm in Processing Plant and The Sixteenth Cavern, although it doesnât finish those.

In Wacky Amoebatrons, Jev completed 9/10 runs. The algorithm completed 1/10 runs. The algorithm almost always ends up at the same spot on the bottom floor, where it shuffles back and forth until itâs trapped by the guardians coming down - by the time it dies, every move is fatal. Jev passes through that spot in some of its runs too, but it moves on. In Amoebatronsâ Revenge, Jev completed 7/10 runs, and the algorithm 1/10, mostly because it got stuck going round in a loop.
So what is Jev doing that the algorithm isnât? The algorithm only looks at two things: does a move collect a key, and is it ânearerâ. Jev also gets facts about the guardians that walk left and right: which side theyâre on, how far away they are, and whether theyâre heading towards Willy. (It gets nothing about the up-and-down ones - we tried that early on, and it made things worse. Maybe information overload?)
To test whether those guardian facts matter, we ran Jev without them. In 18 of the 20 caverns it made no measurable difference - removing the moves that kill Willy is enough. But in the two Amoebatron caverns, Jevâs results roughly halved: 17/20 down to 7/20 complete runs in Wacky Amoebatrons, and 13/20 down to 6/20 in Amoebatronsâ Revenge. So in exactly the caverns where Jev beats the algorithm, itâs the guardian facts that itâs using.
Even without the guardian facts, Jev still did a bit better than the algorithm there, so it gets something from the other facts too - probably the memory facts (âWilly has been here beforeâ, âWilly already tried this move from hereâ), which the simple algorithm ignores. We gave the algorithm the memory facts as well: it now prefers a ânearerâ move that Willy hasnât tried before. That helped it a little in these two caverns, but it got nowhere near Jev.
Two more things we checked:
- Jevâs top choice really matters. Instead of playing Jevâs top choice, we tried picking a move at random, weighted by Jevâs probabilities. That got through only 18% of the game, against 33% for Jevâs top choice (both with Jev choosing its own keys). Jevâs top choice is far better than a weighted coin toss.
- The difference builds up over a run. We took 300 moments where Jev chose a move the algorithm wouldnât have, played Jevâs move and the ruleâs move from the same saved game, and then let the rule play on from both. On average, Jevâs move led to slightly more keys. Each of Jevâs choices is a little better than the algorithmâs.
Fixing the facts
After all that, Claude went back through the failed runs and asked why it failed - was it Jev or our harness. In 122 of the 166 failed runs of one measurement, the facts had let Jev down: either no move was ânearerâ, or the ânearerâ moves went round in a loop.
So we fixed the facts. Instead of guessing a âway upâ from the text map, the harness now explores the cavern in the emulator. It tries every move from every place Willy can reach, with the guardians switched off. âNearerâ now means fewer moves along a route that really exists. Weâre back to having search solve the cavernâŠ
It also turned up something embarrassing. With our six moves, three of the caverns looked impossible. A search found no route through Processing Plant, Return of the Alien Kong Beast or Solar Power Generator, and none of our 1,349 runs of them had ever finished. A walk always ends on the cell grid, and some jumps have to start half a cell over. Two extra âhalf stepâ moves fixed that - with them, the search finds a route through every cavern.
With the true facts, both Jev and the algorithm did far better - and this time the algorithm came out on top. Here, Jev chose its own keys and the algorithm went for the nearest key:
| Share of all the keys and portals | Old facts | True facts |
|---|---|---|
| Jev | 33% | 60% |
| Algorithm | 22% | 71% |
In complete runs, Jev went from 28 to 89 out of 200, and the algorithm from 20 to 114. With the same known-good key order for both, itâs 71% against 55%, and 125 complete runs against 86. Thatâs very unlikely to be luck - about a 1 in 7,000 chance. Between them, theyâve now finished 19 of the 20 caverns.
Two caverns got worse for both of them. The true facts know the route, but not the guardians. In The Sixteenth Cavern, guardians trapped Willy at the left end of the bottom floor in every run, before he got a single key. Solar Power Generator went much the same way with the known-good key order.
Jev still picks a ânearerâ move about 90% of the time when there is one. But now that the facts are right, its other choices cost it runs: it waits just as a guardian comes down, or it goes round in a loop. When several moves are ânearerâ, the algorithm picks one at random, and that does better. Even Jevâs lead in the two Amoebatron caverns shrinks, because the algorithm no longer gets trapped there.
So is Jev any good?
For this game, most of Jevâs edge came from coping with bad facts. Where the harnessâs hints led into a loop or a trap, Jev sometimes found its way out. Once the harness knows the real route, the three-line algorithm beats it.
We also tried Claude Haiku (through the claude command-line tool) with the same facts. It completed Central Cavern in the one run we tried, in 123 decisions, which is not far off what Jev needs (78 to 138 in our runs). But each decision took almost 18 seconds - against Jevâs quarter of a second.
Everything is on GitHub: the harness, the summaries of every measurement in this post, and a viewer that replays the recorded example runs and shows exactly what Jev was given for each decision.
If you see an âAI plays a gameâ demo - with Jev or anything else - swap the model for a dumb algorithm and see what happensâŠ
Moral of the story? Same as last time - you can get a very long way with basic algorithms.