🌈 ESP32-S3 Rainbow: ZX Spectrum Emulator Board! Get it on Crowd Supply →
View All Posts
read
Want to keep up to date with the latest posts and videos? Subscribe to the newsletter
HELP SUPPORT MY WORK: If you're feeling flush then please stop by Patreon Or you can make a one off donation via ko-fi

In my last post Claude and I failed to get reinforcement learning to play Manic Miner, and we ended up solving the whole game with plain old search.

Around the same time, TypeSafe released Jev. Jev is what they call a “System One” model - the name comes from the fast, instinctive kind of thinking, as opposed to slow, deliberate reasoning. It doesn’t generate text, it makes decisions.

You give it some information about the situation (as JSON), a paragraph of instructions, and a list of options, and it tells you how likely it thinks each option is to be the right one. It’s fast (around 270ms per decision for me) and cheap - a complete attempt at a cavern costs about half a cent.

Jev has already been used to play plenty of games. TypeSafe’s launch demos had it playing Doom from structured game state, and playing Wikiracing (get from one Wikipedia page to another using only links). People have since built Minecraft, Subway Surfers and driving demos, and there’s a Minecraft speedrun agent that pairs a planner with Jev.

So: can Jev play Manic Miner? For anyone who missed the last post - you control Miner Willy, and in each of the 20 caverns he has to collect all the keys and then get to the exit, without touching an enemy, falling too far, or running out of air.

This has been a pretty interesting experiment. There is also a big caveat in this work - there are probably better ways to present the information to Jev and get better decisions. More research is required.

It’s also entirely possible that we just aren’t giving Jev good facts to work on. How do you describe a cave from Manic Miner in text?

At first I thought Jev was smashing it - performing much better than I expected.

Then I tested things using a very simple algorithm-based solution - this made me think that maybe Jev wasn’t doing anything clever at all.

But, after a bit more investigation, I came back to the conclusion that Jev was actually doing something useful.

And then we fixed the information we were giving it, and the simple algorithm pulled ahead again.

Jev playing Central Cavern (one of our recorded runs). Willy is the white figure. The yellow box is the key Jev has chosen to go for, and the bars are its probabilities for the top three moves - this is sped up quite a bit

Describing the game to Jev

Jev isn’t multimodal - I can’t just hand it a screenshot and say “what should Willy do?”. It also doesn’t do reasoning or planning. It makes one quick judgement from what’s in front of it.

My first thought was to give it a map of the cavern as text. This is an ASCII map of Central Cavern at the start of a run:

#........A.X....X............B.#
#...............C..............#
#..............................#
#..............................#
#......................XD..X...#
#=============~~~~=~~~~========#
#.............................E#
#===.........GG................#
#............GG..###.X.........#
#====...<<<<<<<<<<<<<<<<<<<<...#
#............................==#
#..............................#
#...........X.......###~~~~~===#
#.WW.===============.........PP#
#.WW.........................PP#
#==============================#

Each character is one cell of the screen. WW is Willy (he’s two cells wide and two high, and so are the guardians and the portal), A to E are the keys, P is the exit portal, G is a guardian, X is a nasty (the poisonous plants and stalactites), = is floor, ~ is crumbling floor, < is the conveyor, # is wall, and . is empty space. Jev gets a legend that explains all of this.

To a person, that’s a pretty good description - you can trace the whole route on it. Jev can’t.

You can see this when we ask Jev which key to go for next. It gets this map, plus a few facts about each key, like which side it’s on and how far away it is. Key D sits on the same top floor as keys A, C and B, and the only way up to that floor is at its far left end. So once Willy has collected E, a good order is A, C, D, B, walking right along the top floor.

Jev's key decision in Central Cavern, straight after collecting E. The key letters are drawn on for this post - Jev only sees them in its text map

Jev picks D, the key that looks nearest - on the facts, D is 5 cells away and A is 20. And it picks D every single time - in all 977 of our recorded runs that got this far, through nearly two weeks of changes to the harness. D gets a probability of about 0.8, and A about 0.04.

The key problem in Central Cavern: Jev's choice after E, and the route that works

Going for D isn’t a disaster - Willy has to go all the way round to the left either way. But Jev didn’t see that. In one run, Willy collected E at decision 14, and it took until decision 67 before he found the way up at the left and got his next key (he picked up A, then C, on the way to D). When asked again at decisions 40 and 65, Jev was even more sure about D (0.92 and 0.90).

Claude (a reasoning model) worked it out from the same map: Willy can only reach the top floor from the far left, so he meets the keys in the order A, C, D, B as he walks right. Following a route across four floors takes several steps of reasoning, and Jev makes one quick judgement.

It’s the same for moves. We tried showing Jev the map as it would look after each possible move, with no hint about which way to go. It finished 1 run of Central Cavern out of 20, and most runs got one key or none.

Reading a route off a map is not Jev’s strong point. We need to describe what is happening - including which way to go.

This description is what our harness is responsible for - all the code that sits between the game and Jev. Claude built it around the emulator and it does quite a bit of work. A slightly worrying amount of work to me.

  • It reads Willy, the guardians, the keys, and the tiles out of the Spectrum’s memory.
  • It tries all six moves (walk left/right, jump left/right/up, and wait) in the emulator. So we know exactly what each move does.
  • It throws away any move that kills Willy or does nothing (i.e. Willy stays in the same place). It also throws away moves into a dead end: for each move, it plays up to four more moves ahead in the emulator to check that Willy can still stay alive - this is where it starts to feel like we are getting into the realm of cheating.
  • It turns everything into facts relative to Willy, including some memory: “Willy has been here before”, “Willy already tried this move from here”.
  • For each move, it adds a fact called progress: does this move take Willy “nearer” to where he needs to go, or “farther” away?

To play the game, we run the cavern, collect all the facts and possible moves and then ask Jev to pick the best move.

{
  "target": {
    "what": "selected key", "side": "right", "horizontal_cells": 16,
    "height": "higher", "way_up": { "side": "right", "horizontal_cells": 6 }
  },
  "to_the_left": { "first_thing": "nasty", "distance_cells": 1 },
  "guardians": [ { "side": "same column", "height": "higher", "direction": "at Willy" } ],
  "progress_measures": "distance to the way up",
  "moves": {
    "walk_right": { "movement": "Willy moves 1 cell to the right",
                    "progress": "nearer", "place": "new place", "tried_from_here": "no" },
    "jump_left":  { "movement": "Willy moves 4 cells to the left",
                    "progress": "farther", "place": "visited before", "tried_from_here": "no" },
    "wait":       { "movement": "Willy stays in the same place",
                    "progress": "same", "place": "visited before", "tried_from_here": "no" }
  },
  "moves_not_offered": {
    "walk_left": "kills Willy: nasty",
    "jump_right": "kills Willy: guardian",
    "jump_up": "no effect: Willy stays in the same place"
  }
}

Along with this comes a paragraph of instructions that explains the goal and what each fact means. Jev’s answer: walk_right 0.71, wait 0.18, jump_left 0.11. It picks walk_right.

Every so often, Jev also decides which key to go for next, as in the Central Cavern example above. progress is then measured towards whichever key it picked.

First impression: Jev is smashing it - this is amazing!

With all that in place, Jev completed Central Cavern in 10 runs out of 10. The Menagerie and a couple of the one-key caverns went well too. It felt pretty good - although, looking back, those are the easier caverns, and most of the caverns were never completed at all.

The map did help Jev’s key choices in a direct test. We asked it for the next key in 13 situations from the first four caverns. With the map and the facts together, it picked the same key as our completed runs in 12 of 13, against 5 of 13 with the facts alone. In Central Cavern, the map made it start with E instead of D. But better key choices didn’t lead to more completed runs. When Claude chose the key order for the first four caverns and Jev did the moves, Jev didn’t complete any more runs than with its own key choices.

Second thoughts: is Jev actually doing anything?

Working with Claude driving Jev was an interesting experience. Claude was able to analyse each run and determine where and why Jev got stuck. And Claude was able to suggest better facts to get it unstuck.

The problem is, every “better” fact moves more of the “intelligence” into the harness.

The biggest one of these is progress. When the key is on the same floor as Willy, “nearer” means nearer to the key. When the key is on a different floor, the harness works out a way up (or down) from the map, and “nearer” means nearer to that spot.

After a while, progress starts to resemble a very obvious hint of the route to take through the cavern. There’s a real danger of solving the cavern using search and then presenting Jev with a bunch of moves along with a massive hint on which one to pick.

Instructions did seem to make a big difference. One of the difficulties here is that Claude does talk like an absolute lunatic sometimes, and it channels this into the instructions:

“If Willy comes back to the same places again and again, the direct way is closed, and he must go a different way, even if that way goes away from the target first.”

Amazingly, this did actually seem to help!

So, we’ve got a real danger of actually just being able to run the harness and then pick the “best” move presented by the harness without needing to use Jev at all.

This basic algorithm works very well:

  1. If a move completes the cavern or collects a key, pick it.
  2. Otherwise pick a move that’s “nearer” according to the harness.
  3. Otherwise pick any move.

This simple set of rules can’t choose the key order, so we gave it a known-good key order for each cavern, and gave Jev the same order for comparison. We ran each cavern 10 times, and measured how far each run got: the share of the cavern’s keys and portal it reached. A complete run is 100%, and a run that collects 3 of 5 keys is 50% (3 of 6).

The first comparison: how far Jev and the three-line rule got in each cavern

Across the whole game, Jev reached 31% of all the keys and portals, and the rule 26%. Both finished Central Cavern every time and The Cold Room most of the time. Jev picked a move the rule could also have picked in 93% of its decisions, against 54% for a random move. Take progress away and Jev almost never finishes a cavern. Random moves barely get started (4%).

Willy's path in Central Cavern with Jev (teal, 112 decisions) and with the three-line rule (orange, 130 decisions). Same key order, almost the same route

So my conclusion at this point was: Jev can play Manic Miner, but the harness is actually doing all the work.

Why do so many caverns stop partway? Most of those runs didn’t die - they got stuck, walking back and forth until 30 moves passed with nothing new. In several caverns Jev got about halfway and then stalled. The harness only knows the way to the next floor, and in caverns where the route winds across several floors, “nearer” leads straight into a loop. It’s highly likely that the “facts” we are presenting Jev work well on the first three caverns, because they are the ones we tested and tuned against.

Digging deeper

After some deep thinking from Claude, we tried some more experiments.

We added a second simple rule, just for keys - “go for the nearest key” - and gave it to both Jev and the algorithm.

Like for like: how far Jev and the three-line rule got in each cavern, over 20 runs each

With the same key choice, Jev now got clearly further than the rule: 39% of all the keys and portals against 25%, and it finished caverns about twice as often. The same happened when both used the nearest-key rule (32% against 19%). That’s very unlikely to be luck - roughly a 1 in 1,000 chance.

On some caverns Jev is better

Jev stands out in two caverns, Wacky Amoebatrons and Amoebatrons’ Revenge: it finishes most runs there, and the algorithm almost never does. It also gets much further than the algorithm in Processing Plant and The Sixteenth Cavern, although it doesn’t finish those.

Wacky Amoebatrons: Jev's routes, and where the algorithm died

In Wacky Amoebatrons, Jev completed 9/10 runs. The algorithm completed 1/10 runs. The algorithm almost always ends up at the same spot on the bottom floor, where it shuffles back and forth until it’s trapped by the guardians coming down - by the time it dies, every move is fatal. Jev passes through that spot in some of its runs too, but it moves on. In Amoebatrons’ Revenge, Jev completed 7/10 runs, and the algorithm 1/10, mostly because it got stuck going round in a loop.

So what is Jev doing that the algorithm isn’t? The algorithm only looks at two things: does a move collect a key, and is it “nearer”. Jev also gets facts about the guardians that walk left and right: which side they’re on, how far away they are, and whether they’re heading towards Willy. (It gets nothing about the up-and-down ones - we tried that early on, and it made things worse. Maybe information overload?)

To test whether those guardian facts matter, we ran Jev without them. In 18 of the 20 caverns it made no measurable difference - removing the moves that kill Willy is enough. But in the two Amoebatron caverns, Jev’s results roughly halved: 17/20 down to 7/20 complete runs in Wacky Amoebatrons, and 13/20 down to 6/20 in Amoebatrons’ Revenge. So in exactly the caverns where Jev beats the algorithm, it’s the guardian facts that it’s using.

Even without the guardian facts, Jev still did a bit better than the algorithm there, so it gets something from the other facts too - probably the memory facts (“Willy has been here before”, “Willy already tried this move from here”), which the simple algorithm ignores. We gave the algorithm the memory facts as well: it now prefers a “nearer” move that Willy hasn’t tried before. That helped it a little in these two caverns, but it got nowhere near Jev.

Two more things we checked:

  • Jev’s top choice really matters. Instead of playing Jev’s top choice, we tried picking a move at random, weighted by Jev’s probabilities. That got through only 18% of the game, against 33% for Jev’s top choice (both with Jev choosing its own keys). Jev’s top choice is far better than a weighted coin toss.
  • The difference builds up over a run. We took 300 moments where Jev chose a move the algorithm wouldn’t have, played Jev’s move and the rule’s move from the same saved game, and then let the rule play on from both. On average, Jev’s move led to slightly more keys. Each of Jev’s choices is a little better than the algorithm’s.

Fixing the facts

After all that, Claude went back through the failed runs and asked why it failed - was it Jev or our harness. In 122 of the 166 failed runs of one measurement, the facts had let Jev down: either no move was “nearer”, or the “nearer” moves went round in a loop.

So we fixed the facts. Instead of guessing a “way up” from the text map, the harness now explores the cavern in the emulator. It tries every move from every place Willy can reach, with the guardians switched off. “Nearer” now means fewer moves along a route that really exists. We’re back to having search solve the cavern


It also turned up something embarrassing. With our six moves, three of the caverns looked impossible. A search found no route through Processing Plant, Return of the Alien Kong Beast or Solar Power Generator, and none of our 1,349 runs of them had ever finished. A walk always ends on the cell grid, and some jumps have to start half a cell over. Two extra “half step” moves fixed that - with them, the search finds a route through every cavern.

With the true facts, both Jev and the algorithm did far better - and this time the algorithm came out on top. Here, Jev chose its own keys and the algorithm went for the nearest key:

Share of all the keys and portals Old facts True facts
Jev 33% 60%
Algorithm 22% 71%

In complete runs, Jev went from 28 to 89 out of 200, and the algorithm from 20 to 114. With the same known-good key order for both, it’s 71% against 55%, and 125 complete runs against 86. That’s very unlikely to be luck - about a 1 in 7,000 chance. Between them, they’ve now finished 19 of the 20 caverns.

True facts, like for like: how far Jev and the three-line rule got in each cavern, over 10 runs each

Two caverns got worse for both of them. The true facts know the route, but not the guardians. In The Sixteenth Cavern, guardians trapped Willy at the left end of the bottom floor in every run, before he got a single key. Solar Power Generator went much the same way with the known-good key order.

Jev still picks a “nearer” move about 90% of the time when there is one. But now that the facts are right, its other choices cost it runs: it waits just as a guardian comes down, or it goes round in a loop. When several moves are “nearer”, the algorithm picks one at random, and that does better. Even Jev’s lead in the two Amoebatron caverns shrinks, because the algorithm no longer gets trapped there.

So is Jev any good?

For this game, most of Jev’s edge came from coping with bad facts. Where the harness’s hints led into a loop or a trap, Jev sometimes found its way out. Once the harness knows the real route, the three-line algorithm beats it.

We also tried Claude Haiku (through the claude command-line tool) with the same facts. It completed Central Cavern in the one run we tried, in 123 decisions, which is not far off what Jev needs (78 to 138 in our runs). But each decision took almost 18 seconds - against Jev’s quarter of a second.

Everything is on GitHub: the harness, the summaries of every measurement in this post, and a viewer that replays the recorded example runs and shows exactly what Jev was given for each decision.

If you see an “AI plays a game” demo - with Jev or anything else - swap the model for a dumb algorithm and see what happens


Moral of the story? Same as last time - you can get a very long way with basic algorithms.

Related Posts

-
Esp32 s3 zx spectrum - In a bid to quench my nostalgia and flex my ESP32 chops, I managed to get a ZX Spectrum emulator running on my ESP32-TV board! Then, spurred on by PCBWay's new full color silk screen service, I pursuit the audacious task of recreating the ZX Spectrum's iconic keyboard. It's been quite the joyride - wrangling touch pins, shrinking screens and creating a thing of beauty on PCB. It's not quite ready for the spotlight, but keep an eye on my newsletter for more eagerly-awaited updates. It's like the Spectrum is reborn!
HELP SUPPORT MY WORK: If you're feeling flush then please stop by Patreon Or you can make a one off donation via ko-fi
Want to keep up to date with the latest posts and videos? Subscribe to the newsletter
Blog Logo

Chris Greening


Published

> Image

atomic14

A collection of slightly mad projects, instructive/educational videos, and generally interesting stuff. Building projects around the Arduino and ESP32 platforms - we'll be exploring AI, Computer Vision, Audio, 3D Printing - it may get a bit eclectic...

View All Posts