Launcher

Type to filter results. Use arrow keys to navigate, Enter to select.

You've reached the end of the results

Every time I come back from a big pinball event where I’ve had some scorekeeping duty, I find myself thinking about better ways of doing it.

Tournament software has come a long way. Tools like Matchplay and DTM are excellent, and they’ve made running large tournaments dramatically easier than it used to be. Pairings, standings, queues and results all just happen, and players can follow along on their phones.

Some parts of scorekeeping are genuinely fun. At a big tournament you get to meet loads of players, put faces to names, and have dozens of tiny interactions with people as they come through.

But then there’s the other bit.

Type a number. Next game. Type a number. Next game. Type a number.

Repeat for an hour or more and your brain begins to feel mushy.

So it got me wondering whether there’s still some room to optimise the data entry side of the process.

The Concept at a Glance
  • The Problem: Scorekeeping at big pinball tournaments is a mix of the fun bit (meeting players) and the dull bit (typing in score after score).
  • The Idea: Take a photo of the machine instead. Software reads the score off the display and works out which game it is, so the score lands in the right place.
  • Reading the Score: AI vision models turned out to be surprisingly good at this.
  • Knowing Which Game It Is: The harder bit. My first attempt, automatically matching the photo against pictures of known games, was unreliable and hard to debug.
  • Where Jev Comes In: A vision model (Qwen-VL) describes what it sees in plain English, and Jev picks which of the tournament's games best matches that description.

What If the Camera Did It? (Testing Qwen-VL & OCR)

There are already robust solutions for capturing scores automatically. Scorbit, for example, takes the idea much further, with hardware fitted inside the machines that reads the score straight from the game. That’s a very different approach, and one that has obvious advantages when you control the environment.

But it did make me wonder about something much more basic. Pretty much every phone and tablet already has a camera. Computer vision has improved enormously, and modern vision-language models (AI models you can show a picture and ask questions about) are getting remarkably good at looking at pictures and describing what they see.

So what happens if you just point a camera at a pinball machine?

After scorekeeping at the 2025 UK Open, I started playing with this as a side project. I wanted to know if a scorekeeper could snap a photo of the display instead of typing in numbers.

To test it, I wired up a Telegram bot backed by an n8n workflow, feeding photos into Qwen3-VL (which was my go-to vision-language model at the time, and has since been superseded).

That combination turned out to be great for prototyping. n8n handled the pipeline behind the scenes, wiring up the API calls, parsing the JSON and holding the prompt, so I could tweak it all in one place without touching any app code. The prompt needed a fair bit of tweaking too. Pinball displays are cluttered beasts, and you have to tell the model where the player scores sit relative to credits, ball counts or scrolling match animations.

Telegram was brilliant for fast, real-world testing. Having the bot on my phone meant there was no app to build. I could stand in front of a machine, snap a photo, and get an answer back in a couple of seconds. Then snap another from a steep angle, or with reflections on the glass, or from further back, and immediately see what the model made of it.

Once the prompt explained how score displays are laid out, it was surprisingly resilient, and far more reliable than the dedicated OCR libraries I’d wrestled with before.

Telegram bot using Qwen3-VL OCR to read player scores from a White Water pinball display showing 22,670,780 and 1,100,110
Reading scores from a photo. A two-player White Water game sent to the Telegram bot. The n8n workflow passes the photo to Qwen3-VL and replies with both player scores a few seconds later.

The Real Challenge: Identifying Pinball Machines Visually

In an ideal scorekeeping workflow, the scorekeeper shouldn’t have to select the game or type numbers at all. In tournament platforms like Matchplay, the software already tracks player queues and pairings. The scorekeeper just needs to walk up, point the camera, and confirm.

Take a picture → Identify the game → Read the scores → Done.

If the camera is doing the work, the interface gets a lot simpler. The scorekeeper is only there to check what the machine vision came back with. That probably doesn’t need a tablet at all. A phone would do.

Reading digits off a display is a well-trodden problem with modern OCR, and facial recognition is everywhere. Loads of engineers have poured years into both. But telling one pinball machine apart from another? That’s not exactly a common challenge, is it? It’s a niche, visual problem, and for me that was the really interesting bit.

Why Image Embeddings Hit a Wall

My first attempt at game identification was the fairly standard approach. I started looking at vector embeddings, which turn an image into a long list of numbers so that similar-looking images end up close together. I pulled some data from Pindigo (now becoming Flippd) for a small selection of games and tried to build embeddings that could match a photographed machine against known games.

Telegram chat: vector embedding match for Blackwater 100 pinball machine with a 0.26 similarity score
Matching a game with embeddings. A photo of Blackwater 100 sent to the bot. The top match is the right game, pointing back to the Pindigo photo it matched against, but with a similarity score of only 0.26.
Telegram chat: vector embedding mistakenly matching Blackwater 100 as Jack*Bot with a 0.29 similarity score
Same game, different answer. The same Blackwater 100 from a slightly different angle and crop. This time the top match is Jack*Bot, with a higher score than the correct match above.

And, occasionally, it worked. But take another photo of the same machine from a slightly different angle or with a tighter crop, and it would suddenly decide it was Jack*Bot.

That was even when choosing from a very small subset of games. When it did get the right answer, the similarity score was low. And I had no real way to see which features of the game it was using to make the match. Was it the backglass artwork, reflections on the glass, or the lighting in the room? Making it robust was going to need a lot of data, with photos of every game from different angles and in different lighting. It was rapidly turning into a research black hole.

Of course, this is exactly the kind of technology that’s interesting to Flippd and Pindigo too. It’s the same problem set. A player photographs the game, the app needs to work out which machine it is, and then someone has to type in the score. So I reached out to the Flippd folks on their Discord and had some interesting conversations with people working on similar problems from different angles. At that point I had one of those moments where you look at the project you’ve accidentally created and realise it has absolutely nothing to do with the physical pinball platform being built at MAYA Pinball.

So I stopped. This was supposed to be a fun side experiment, but the scope was creeping.

Enter Jev: Constrained Classification Without Embeddings

Fast forward about a year. I’m at EPC and the UK Open, doing scorekeeping on my tablet, and exactly the same thoughts come back into my head. At the same time my technology feeds are full of people getting excited about Jev.

Jev is a new kind of AI model from TypeSafe AI. Unlike a chatbot, it doesn’t write text back to you. You give it some information, you tell it the set of possible answers up front, and it picks one, along with a probability for how confident it is. It’s built for quick, discrete decisions rather than open-ended conversation, and I wondered whether it might change the shape of this problem. In true computer science fashion, jev-fit is a fun, recursive way of describing what Jev is. You paste in an idea and Jev decides whether it’s a job for Jev.

What struck me is that it gives you a completely different way of thinking about identifying a game. Instead of turning the photo into a list of numbers and hoping the closest match is right, what if you just describe it?

You could describe the artwork on the backbox, the dominant colours, the characters, the logos, the layout of the playfield, distinctive graphical elements, or anything else visible in the photograph that helps distinguish it from another game. Modern vision-language models are remarkably good at exactly this sort of thing.

The interesting bit is that this feels much closer to how a human actually identifies a game. There are games I’ve seen hundreds of times where I can probably tell you what they are before my conscious brain has even finished processing the image. But even for games I’m less familiar with, I can often tell apart machines from the same manufacturer by looking at the artwork, fonts, and visual cues.

So perhaps the decision doesn’t need thousands of example photos of every game. Perhaps each game just needs a rich natural-language description of what makes it visually identifiable.

Why Not Just Ask the Vision Model to Name the Machine?

At this point you might reasonably ask why not skip Jev altogether. If modern vision models are so smart, why not just ask Qwen-VL “What pinball machine is this and what is the score?” in one prompt?

Maybe that would work well enough for well-known games. But a chat model answers in free text, so you’d still need to match whatever it writes back to a game in the tournament, and you don’t get a clear sense of how sure it is. Jev picks from a fixed list and tells you how confident it is.

So the idea is to split the job. The vision model does the observing, and Jev makes the decision.

Jev doesn’t take images directly (at least not yet), so the vision model still does the looking. But instead of asking it to guess the machine name, you ask it to read the score and describe what it can see:

{
  "score": 9802471,
  "observations": [
    "Bally logo on the backglass",
    "motocross riders on dirt bikes",
    "large 'Blackwater 100' title in yellow and black",
    "snake artwork in the bottom left",
    "seven-segment score displays"
  ]
}

Then Jev gets those observations, plus the list of machines it’s allowed to choose from, each with a short description of what makes it recognisable:

blackwater_100:
  - Bally logo on the backglass
  - motocross / dirt bike riders
  - "World's Toughest Race" on the backglass
  - muddy brown and green dirt track playfield

jack_bot:
  - Williams logo
  - large robot character
  - casino and space theme
  - dot matrix display

Which machine is most consistent with what the camera saw? Jev picks one, with a probability attached.

Identifying a machine with a vision model and JevA photo goes to a vision model, which reads the score and lists what it can see. Those observations go to Jev along with the list of candidate machines. Jev picks the machine. A new machine is added by adding its description to the candidate list.Photo from phoneVision model (Qwen-VL)reads the score, lists what it seesJevobservations + candidate listBlackwater 100machine + score, with a probabilitynew machine?add its description
Identifying the game. The vision model describes the photo, and Jev picks from the machines in the list. Adding a new game just means adding its description to that list.

The really nice thing about this is that adding a machine is just a case of writing a description for it. Start with five games in the list and Jev only has to tell those five apart. Add another five later and you haven’t retrained anything, generated embeddings or tuned a similarity threshold. You’ve just expanded the choices.

It’s also easier to reason about when it goes wrong. Remember the embedding results? One photo correctly matched Blackwater 100, but with a score of only 0.26. Another photo of the same machine matched Jack*Bot with a higher score of 0.29, and was completely wrong. With embeddings, you’re often left asking why. Here, everything going into the decision is plain English. You can read what the vision model said it saw and read the descriptions it was compared against. If it keeps confusing two games, you can see which observations overlap and improve the descriptions.

I’m deliberately hand-waving the actual implementation here. I haven’t built this part. I’m just wondering if this is a more human-shaped solution to the problem.

The Tournament Advantage: A Short List of Games

Since I originally played with this, Flippd has opened its API again, which makes the whole thing considerably more interesting.

You could potentially take their catalogue of games and have a vision model produce observations about the visual characteristics that make each one distinctive, distilling those into the descriptions Jev needs.

There’s a trade-off here. If you generate highly specific descriptions using one particular vision model, and then build the system around those descriptions, you may end up tightly coupling the data to the model that produced it. Swap the model later and the descriptions might not line up so well. On the other hand, perhaps that’s completely irrelevant. Plain-English features like ‘motocross riders on a dirt track’ or ‘Williams logo’ should mean the same thing to any model.

More importantly, tournaments make this much easier, because you don’t have to identify every pinball machine ever made.

You only need to identify the machines being used at that specific event. If there are 30, 50, or even 100 games in the tournament bank, you’re not asking the system to pick one machine out of the entire history of pinball. You’re asking Jev to choose from a short, known list, which is exactly the kind of decision it’s built for.

Beyond Data Entry: Livestream Overlays and Next Steps

The original experiment was shelved because building and tuning vector embeddings was rapidly turning into a full-time machine learning research project. Pairing modern vision models with Jev might change that. On paper at least, it starts to look more like a prompt and a list of descriptions than an ML pipeline.

If it works, the scorekeeper’s job gets a lot lighter. No manual game selection dropdowns. No typing 10 digit numbers. No add-on hardware inside every cabinet. Just the camera on the device the scorekeeper is already holding.

And the applications don’t stop at tournament scorekeeping. Plenty of tournaments and players livestream their games. The same pipeline could pull live score updates and game titles directly off a stream’s camera feed to drive overlays, graphics, and live leaderboards, without anything plugged into the machine.

Take that one step further and you don’t need anyone to point a phone at all. Put a high-resolution camera over the whole bank of games and let it watch. It sees which game each player is on and reads the score when the game ends. You’d end up with something like the Amazon Go shops, where you take items off the shelves and just walk out. Play your game, walk away, and the score is already in.

What Do You Think?

There are undoubtedly real-world edge cases to iron out. Honestly, I’m mostly writing this to get the idea out of my head rather than disappearing down the rabbit hole again.

I also think this is only scratching the surface of what Jev could do in the context of pinball. People already have it playing video games like Doom, making split-second decisions in real time, so identifying a game from a description feels like a gentle place to start. If it can play Doom, we might have a whole new type of pinball attract mode in the not too distant future!

Have you tackled pinball score OCR or game identification, or do you have thoughts on the Jev architecture? Get in touch, I’d love to hear your take.