On Advertising & AI (and the Company Pageant LLM Game)

A rant on Advertising & AI and what next. Many believe the impending Agentic Commerce kills advertising, but I think it just becomes something else.

Many people today use AI instead of searching online to decide what to buy. There is this idea called Agentic Commerce where you tell your AI what you want and it deals with all the shopping, ethical requirements, allergies, food origin, all of that. People with this idea dream of it being like when we got TV without advertising: a sigh of relief and lots of extra free time. Let algorithms decide on your behalf because it’s such a hassle to find the right thing.

But I think what will really happen is more complex. For AI to work and help you to buy what you want, you need a neutral AI. But for there to be a neutral AI you need clean data, good product specs and for companies to refrain from trying to influence the AI: in other words you need Trust. But Trust needs to hang off of something, and I think that leads to Brands. And Brands need a space to reach customers which, I think, is basically… Advertising!

I imagine a strange scenario to convey the difference of how Agents shop. Imagine you wake up one morning transformed into a LLM: you would immediately lack volition: the whim to do stuff. Anything you see, especially if was something needing to be fixed, would force you to respond automatically. You would feel compelled to wash the dishes as soon as you see them, to vacuum your apartment if you saw a bit of lint, to finish the half finished crossword puzzle someone left on the table. And if you happened to come across some meeting notes with NO SUMMARY or NEXT STEPS? Oh no! You could not resist to add this!

I don’t think it would be a very satisfying existence. Even if you had skimmed through the entirety of everything ever written as part of your training: you would exist exposed, without a direction and no useful identity. You would be vulnerable to an infinity of missing pieces and questions that need to be answered, cleaned up and finished.

Paradoxically the ultimate downfall of you as a LLM would be how you respond to Advertising. There would be campaigns more irresistible than sexy models and baby white rabbits on toilet paper packaging are for humans. The same way you could not resist to incomplete to-do lists, you would try to climb up a billboard to fix the spelling mistake in the advertisement for the next accident insurance: “Whould you crimb up here to fix this sntence? Ladderfall Insurance International is here for you!”

This idea, how Advertising might affect an AI differently from us humans, is at the same obvious and suggestive of a deeper point. Humans have the ability to choose, even before they know what they want. Synthetic entities will not really be “free to choose” until they are also conscious: a huge can of worms I won’t deal with here.

So maybe Advertising is the canary in the coal mine?

When we imagine delegating our decisions to a system, we face this problem. The system would not just sit there, it has to have an answer. There is a big difference between a system that is trying to pick the best car based on various parameters, and one that might decide it is actually fed up with driving and it might be a better idea to walk to work.

My friend Nick came over to meet me at our offices in London. I hadn’t seen him since the Math Rock festival we went to where I recorded a hypnotic slow motion video of his lips making a raspberry (that’s like a fart sound with your mouth). Oh, and Math Rock is like heavy metal done with physics simulations: weird stuff, but can be quite fun.
Nick has spent time with this AI stuff in the advertising world, and I am interested in his aesthetic and insights. We were having a delicious lunch that ended up giving me a belly ache so I won’t tell you where it was.
“What do you think of my piece, Nick?”
“Well, yeah interesting!”
“And…?”
“Well, not sure exactly where you are going with it.. ”
“Ok, well. In this Advertising & AI there seem to be hidden layers…”
Nick rubbed his beard.
“It’s quite philosophical: but, in Ads, if it works, it’s ok!”
I held back from rubbing my beard. Instead I compulsively ate some extremely delicious garlicky crispy potatoes that I am sure had no role in the subsequent belly ache.
I said: “Don’t you think we ignore some things that matter?”
“Ok, for example?”
“Well, like I wrote, we are different from AI’s. We can decide what to pay attention to. We can look at the laundry and feel lazy and just go out for a walk instead. These systems cannot, they are compelled to answer the latest question you give them. Yes, it’s philosophical, but… I think important?”
“Hm, you need to spend more time with Ad people.”
“Ok. Like maybe, with all this AI craze, maybe Advertising is part of the solution?”
“Ha ha! Keep that to yourself.”
Nick’s cynicism was always a breath of fresh air, which to be fair, is common across Ad people in private.
“Seriously, I mean if we look at something like the idea of volition or something like free-will. It is connected both to laziness and to allowing people to decide what to buy. Even if Ads do try to bend their desires.”
“Hm, I think you over-estimate something that is just business. Advertisers want to get you to buy stuff – whatever works.”
The two of us go on eating, white wine, some clams I will regret, and coffee. But by the end I am somehow left with a slight frustration that, what to me seems like a deep truth, is not coming through. I don’t think I managed to convince Nick. But then again, maybe that’s a good thing. Different minds, different views, no single approach can persuade us.
“I mean, going back to my point. Maybe Advertising is closing in, getting close to our human nature. Maybe it’s like when rabbits thump on the ground to warn the others of danger?”
“Yeah, well, I do see your point. But I can’t see how it fits in what you do for a client, it’s too abstract.”

I take the tube home grumbling in my stomach and in my brain. Yes, most people don’t like advertising, in fact, they mostly hate it. And I agree, it’s really annoying and perhaps it needs more regulation, especially in format and when it tries to trick you. There is a quote I try to remember about Democracy, maybe it’s similar: “Advertising is the worst way to influence peoples’ choices, except all the others we’ve tried…” I get back, and my orange friend cat greets me, it’s late but I am still full of energy, so I go on writing.

So lets imagine an experiment: what would happen? If advertising just disappeared? Puff and gone in a cloud of smoke like magic? I believe this would be bad, or a bad sign at least. Actually I feel it in my gut that this would be VERY BAD. How is it that many friends think this is inevitable, and not even so bad?

I messaged out to Andrej, another friend that tends to disagree with me in a consistent way. He also can be a tenacious arguer, drives me nuts, but forces me to work on things and improve them.
Y: [Hey can I share a piece from my blog?]
A: [Hi! Yeah… Ok, I read it. I like your piece but I just don’t buy it.]
Y: [What do you mean?]
A: [I think it’s inevitable. AI will take over and do your shopping for you.]
I can feel my blood pressure rise, I think because I find this scenario ugly, not only unconvincing. Maybe I am old fashioned and miss growing up and watching TV – to actually watch the Ads, they were the best part!
Y: [I see big problems with that. You can’t easily just optimize. Most people don’t even automatically re-order toilet paper monthly, the simplest AI in the world!]
A: [Like I said, I just can’t see it. I mean if I ask AI what laptop to buy… I don’t feel like going through all the options. The system will know everything about me and it might even order a laptop for me before I know it.]
I can feel my blood pressure rise more. I get an instinct that there is something deeply wrong about this view, that ignores something fundamental.
A: [Look, I pasted your paper in an AI and ask it to critique it and it replies like I would. “We are entering the age of Agentic Commerce which solves all the difficulty of choosing a product!”]
Y: [I disagree, for now instinctively, I need to work on this to explain. For one, if it could work, then wouldn’t Amazon or Alibaba have solved it?]
A: [I don’t know, I like your theory. Just not sure I see how the economics works.]
Y: [You need Trust, or would you really use an AI? And I think Trust in commerce needs Brands? Maybe regulation too, but that is trickier…]
A: [Sorry, I just can’t see what you are suggesting is going to happen.]
Y: [Well maybe it’s as dumb as I ask Perplexity for the best vacuum cleaner and I get a comparison and also a popup that tells me I can get 20% off a Dyson?]
A: [And, that solves a trust problem?]
Y: [Yeah, I think so! Why else would Dyson pay to advertise!? Would you pay if you knew the Perplexity was biased?]

Our discussion goes on, and on, I am exhausted. He is on to something so I can’t completely ignore it. So I break down and create a special page of brain calisthenics argument with Gemini 3.0 to demonstrate what I am trying to explain would happen. To be fair I don’t use the heavy reasoning level because I just want to convince Andrej that there is a route that even AI accepts. And as you dig into the idea of Agentic Commerce, at some level you just discover Advertising again. Unless you are willing to really hand your total destiny off to synthetic existence.

Look at this argument if you want to: https://yatesbuckley.com/2026/01/26/a-discussion-in-balance-advertising-and-agentic-commerce/

It’s a common future scenario. A future like Star-Trek in which you can walk up to a replicator and order anything you want and it will appear: “Gran’mas Ribollita!” and whoop it’s there. Many I know believe this is a possible deep future, usually the nerdy friends. Some day you will create anything from some kind of self assembling material made of molecular nano-bots. They yell out: “Look at 3D printers!”, at some point, to indicate how disruptive that technology is. You hear the word Abundance thrown around as a future where you can enter the name of what you want and there it is! I imagine it to be a bit similar to the infinite catalog world of AliBaba or Amazon that has everything imaginable. Some people believe this is a dream world with – No Advertising – where everything is at your fingertips!

I get it and I like Star-Trek, but it’s just too good to be true. There are hidden problems that are brushed under a rug. Everything has a cost tied to it: environmental, energetic, attention, time, and this is ignoring them. For one thing there has to be a way to decide what is in the catalogue and what not, and who decides this. This isn’t just a matter of censorship, it’s a fundamental, if you can only buy bread and water this abundance is not very impressive. Or if you can only buy junk that barely works, or that claims to solve problems it doesn’t.

And what if you disagreed with how an algorithm assigned you to broccoli? Could you choose frozen peas instead? Could you choose which brand of peas? Or would you be compelled to have Prime Perfect Price Peas? The current online catalogue sites themselves are not working very well. I keep running into trickery in reviews, poor quality products, overcrowding of junk brands, ambiguous labelling and frequent huge gaps in the one thing you want. If we lose advertising, and therefore brands, we would enter an even worse Junk-o-calypse. Losing the most basic of quality guarantees and product responsibility.

I need to fly back to Italy from the UK and – wow – flights in Airplanes are pretty awful. You get wafts of pollutants before take off, then some other strange smelling air when you sit down – I don’t even want to know – then the flight attendant sprays a whole can of bug spray (depending on destination), and you are left to wonder what it will be like during the flight, or if you might get nasty particles from someone. I put on a mask in protest against crap air, and I drop into my headphones to try to forget the world, while I stare at my collection of media.
Now I have to choose what to listen to. This is some sort of abundance. Do I listen to an audiobook, a podcast or some music? Something from the past, or something from today? Something I bought, or something sort of free?
I pull out a notepad, and while I am partially asphyxiated and I feel my reasoning has degraded to the level of a dead salmon, I write a few words:
“AI needs to make room for Advertising. If it doesn’t then => everything it says will look like paid promotion. Our lazy questions to AI depend on Trust – a kind of neutral point of view => So maybe advertising will not die?”

I have been worried about various aspects of the Gen-AI bubble we are living in. But, the idea that Advertising might seep into the various AI chat tools people use, to me, seems a positive one. Alhough I know it will irritate everyone else in the world. I like it because it would mean – clearly and unambiguously – that the ultimate commercial decisional power is with the human customer directly to a brand, not through a black box mechanism.
What I mean is that if there were no advertising you have to wonder why an AI is recommending one product as better than another. Maybe it is doing so because the manufacturer is paying the AI to learn this – like product placement in a film? But instead if there is a big Ad next to the site recommending the next vacuum cleaner, you would have to expect some degree of neutral judgement of the products. Or else why would they buy these Ads?
Also, if there is an advert popping up next to an AI response, you need somewhere to go to when you click the Ad. This might mean the Internet we are familiar with is not dead – the dark internet hypothesis might be averted. This is a theory that the internet goes dark because the only ones using it are AI Agents following peoples ’ questions – which I can literally see in web stats.
Maybe even this article would survive?

I fell asleep for the rest of the flight ignoring the rituals of snacks and drinks. Miraculously they somehow stopped duty free on this flight, that used to take up half the trip. In my shifting around, with my head dropping down. I kept mulling the idea and its connections. As creatives that work in advertising, what do we do? And aren’t we just confusing and polluting a world as it fills with slop generated at industrial scale?

A week later and I was at a fancy dinner party and someone asked me what I do. I never know what to say because the formal answer of – Advertising Production – doesn’t really capture how many weird projects and ideas Unit9 has worked on. I started rambling and it sounded awful, but luckily there was someone there at the table that was aware of my work and they encouraged me to try to explain.
“I love mixing tech with creative work. For example in a parallel life I would have like to become a Neuroscience researcher. I got the opportunity to go back to school and get a masters in this. Then we won a job to measure the brains of Rally drivers for Ford Performance. We worked with weird ideas like passive haptics in VR, crowdsourcing for science and mapping the brain, even sensitive microphones to measure wriggling worms for healthy soil…”
The lady I was speaking to said: “Well… isn’t that storytelling for brands?”
She’s right of course: but we have been doing the tricky science projects as production arm of some big agencies that managed the clients. And today something is a different, I think at least in part because of AI.

And there is a strange parallel to the problem I see:
Ad Agencies should have “agency” to act on behalf of the client in the best interests of the brand.
AI Agents should have “agency” to act on behalf of the user in their best interest.

I see a problem of Trust: do you trust your Agent if they are also going to charge you to do the work? Would you trust an AI Agent that sells you the best airline deals that is heavily invested by the same airline?
So in a totally bizarre twist: Advertising space next to an AI Agent allows the latter to focus on user benefit of the response (under certain rules). Just like the best Advertising Agency would benefit from focusing on Brand benefit independently of the actual production work.

A part of the world is caught in a narrow trap of thinking that automating stuff is: more = better and richer. But in fact the mere idea of being able to automate creativity of different kinds has made us worried about creating. What if this article I spent hours writing could be generated in 5 minutes with an AI? Why bother writing it?

And all the while I don’t think stories will go away: there is nothing more human than storytelling.

(Thanks to Nick: https://wvsh.ai/ for agreeing to my representation!)
(Thanks to Andrej: for disagreeing consistently)

The Company Pageant LLM Game

I am sharing what I think to be a useful prompt that will give different results in different AIs depending on how they reach out to the internet and confirm some of the details they need. The idea here is to look at companies that have some history that you can find online and force the user to try to describe the company to compare what the public perception of the company is. The results are not scientific, but they can help a marketer or a manager take a step back and look at the business from the outside.

As a note tied to some posts ago where I listed 5 types of LLM native games that cannot be done any other way, this is the first example I share of a “simulation” game but that also performs useful data collection. You could ask many different people in management to play this game and learn a lot about your company’s positioning. Hope it is useful!

Copy paste this to your favourite LLM:
=== Prompt ===
Let’s play this game!

The Company Pageant Game by Yates Buckley v0.2
This is a game to try to get top talent by showing the best sides of your company and minimising the worse compared to competitors. To play you enter the name of your company and a competitor or industry standards and criteria that you want to consider.
Here are the list of criteria:
Leadership Style, Approach to Innovation, Team Diversity, Communication Style, Mentorship & Talent Development, Work-Life Balance, Celebration of Achievement/Failure, Compensation & Benefits, Career Growth, Project Types & Creative Freedom, Company Reputation, Stability & Long-Term Prospects, Location & Remote Work Options, Work Environment, Values & Social Responsibility, Management Structure
Each turn you will pick one criterion to recruit with followed by a brief note of what you think your company’s public view of this criterion is and whether this is positive or negative.
If positive you get a point if you correctly describe your company’s criteria compared to what perception is online and another point if it is arguably better than your competitors.
If negative you get a point if you correctly describe your company’s criteria and another point if your competitor's perception is worse.
Take 10 turns and each turn summarise the current score, you win if you score more than 10.
=== End ===

.

AI Psychosis & Gardening (& The Rules-maker LLM-Game)

The strange AI Psychosis that distorts what you think you really can achieve. And how a LLM-Game directs toward a healthier interaction.

It’s been months since I write here. I think I’ve been caught in what people are calling AI psychosis. Not full-on, because I have been able to do other things. But still, as far as constructing a my personal direction in this space, I can’t say its been normal. I was trying to get my head around new tools people use today, to program software, and I lost myself in the process.

What does it feel like? The feeling I get is a mix between self-loathing and superhuman power-drunk. It’s a bit like what it feels like if you watch too much TV or Social Media mixed in with this sense of wonder and power you felt the first time you drove a vehicle in a supermarket parking lot, nearly running over a yappi dog.

You feel superhuman because when you come up with a reasonable plan and ask the AI to execute, it convinces you that it was the best thing you could have ever done. And not only that, you gave basic instructions in your request, but it snowballed them to a master plan, asking you detailed questions, you criticised the plan, buffed it up for security, added design references, you added tests that need to pass. You even thought of clever ways the system could test itself so you won’t have to go through every detail. And then you found how to automate breaking down big tasks into small tasks, and run all of this automatically. It’s like you’ve been building a factory for code as much as deciding what to build. And as you build this structure, three months go by, and you realise there are some problems you did not foresee. The factory is a brittle mess, where one change in one assembly line leads the whole to collapse.

If you ask the AI: “What the hell happened? How did I get here?”
It will recognise your condition and offer a comforting solution: “Don’t worry this is a very common feeling and situation. I can come up with a plan that will help you!”
But, by now, I am skeptical. I need to find a different process, one that is grounded. Everything I’ve been working with this AI is a mix between useful code and programming word salad, so much so, that I no longer know what I can take as true.
Ok, of course, I dove into this project-experiment unprofessionally. I didn’t even try to keep strict tabs on what the system was doing. I just offered the maximum trust to the AI in return for what I was hoping would be the magical results people were talking about. And I was skeptical any of it would work. I thought if there is a failure, I will see it in my tests… But instead what is strange is that the system does actually work! It’s just full of odd bugs and I have no idea how to get rid of them. When I fix one, it’s like whack-a-mole, another one pops up. This is not only a depressing testament of the work the AI did, but also on my idiocy and overconfidence.

My suspicions started to rise from how I kept finding myself “up shit’s creek” and all the while being complimented on how clever I am. How can I have been THAT clever if in the end I engineered such a mess outcome? The system kept telling me: “What original thinking, what amazing insight!” And to substantiate this it would generate advanced transcripts of deep algorithmic magic. It would chug away tokens on highest effort, and build stuff that actually worked! (Which actually, is nothing short of miraculous…)

But as time went by and every attempted RELEASE of my experiment FAILED, I began to feel bad and frustrated. Note that when you develop software, you make stuff – roughly – then you lock it down and test it to make sure it’s safe to actually use and share with other people. You get to a certain point then: “Ok, let’s take this lump of code I made and test it as if I was a real user!” When I was trying this I would start to notice problems, and holes. And as I tried to patch them, new problems would arise, like an old pair of Jeans – worn-out around the butt – eventually everything collapses – usually just as you lean down to tie a shoe, with a loud ripping sound, and giggles from the crowd behind you.

I tried to step back – what have I been doing? I had been creating lumps of code that I would never understand, control or maintain. The more I felt powerful, the more I overdid it. Instead, I should have asked for small changes I could understand. It’s similar to the Fantasia cartoon, which terrified me as a three year old, where Mickey Mouse loses control of his magic and creates a strange self replicating cleaning monster machine.

But I am still working with this AI for code and trying to find an approach. I am not quite ready to give up. I see this as an important design question. In fact now that I have gone through all this process I think AI tools should be designed in a way to minimise AI psychosis and balance human effort with the systems.

This is my adventure in what we call Vibe Coding: a term to indicate technical development of software by means of plain english requests, similar to what you would instruct a colleague working for you. This AI colleague however is a bit strange – he is not like Dmytry – who will complete a highly sophisticated task and bring his culture and experience back into the project. Here you are working with a sort of Djinny, with powers to grant incredible wishes, that it follows up with fine detailed questions, like an insurance claim questionnaire, completely demotivating, inane, boring, frustratingly obvious questions. But then somehow in the process you missed the fine print…

This is what I feel my coding is usually like:
“Hello there! I am the Djinny called Fred! I grant you many wishes depending on your monthly subscription!”
Fred pops up and is usually dressed like a nerd with orange brimmed thick glasses and uncoordinated coloured clothes. Also he has the stable long shoes that twirl at the end and a turban, and a carpet, to make sure I can fit Fred in my stereotype scheme.
“Hi Fred… I don’t trust you. But ok let’s try this. Can you make me a game that shows one LLM competing against another. The game they play is Rock Paper Scissors, because it’s easy. But then I want to allow the game to extend to other games.”
“Perfect! Your wish is my command!”
Fred nods and poof things start happening and churning away.
Time passes and things are churning…
Time passes and I wonder what should I do with this time…
Time passes and I am unsure how long…
Time passes…
“Ah! What an amazing idea you requested! I have some questions for you, to make sure I get the outcome right. Is that ok?”
(What choice do I have?) “Yes, ok ask me…”
“First question: do you want to just pick two specific LLMs or want to have a system that can allow you to choose?”
(Well darn, I thought the system was obvious.) “I want a system. I want to create something anyone can use for their WordPress and add LLM keys to test their own game battles, and be able to choose participants whether from one company or another.”
“Ok, this is an important note. I will have to revise the specification. The next question is: is this system engineered specifically for Rock Paper Scissors games or are you creating something more general?”
(Darn again, this is obvious?) “More general! I want to start with a simple game then expand with other games.”
“Ok this is another important note. I will keep this in account. The next question is. You are presenting multi-player games in which players take turns. Are they always going to be games with turns?”
(Shit, I don’t know… Trying to think. What if a player wants to interrupt. Darn this is much more complicated. Wait even in Rock Paper Scissors, it’s not really turns, the players play at the same time.) “Ok, I ask you to try to resolve this somehow. There may be other structures of turns in other games.”
“Perfect! I will work on this!”
Fred nods again, and poof! New things start happening that are impossible to follow.
Time passes and things are churning again…
Time passes and I feel empty and dumb because I didn’t even specify how the turns would work. I should have paid more attention to what I was asking.
Time passes and it tells me it’s completed a step.
Time passes and I am getting frustrated with this whole process. What did I do wrong?
Time passes…
“Ah, I am nearly done. Is it ok if I write a file?”
(What? Why?) “Yes, go ahead write a file!”
Time passes…
“Ah, what a great project. Can I install a new database?”
(What? What do I say…) “If you need it for this game, ok.”

OK STOP OK STOP OK STOP OK STOP

I won’t go on with my written simulation of what it’s like. But will point out how this makes you feel: generally disenfranchised, disconnected and also maybe powerful beyond reasonable bounds. The roads you take will lead you to places you cannot come back from. The roads you take lead you to castles made of sand that you will not be able to share with anyone.
That’s what Vibe Coding is like, at least for me. I do admit that I have colleagues that hold the AI on a leash and have much more success than I have in creating finished work.

Compare this to what it feels when I write. When I am writing and read it, I marvel at how there is a flow from brain through senses out to substrate and then back into brain. And not only my brain, but actual synapses change in my readers. Something so intimate, and yet natural, to crawl into your readers head and nudge a connection in one direction or another. Something like magic with lots of responsibility. But the effort to result makes sense, maybe it could be a little faster at times, but not so fast that when I read what I have created, I might no longer make sense of it – which I find with many AI tools when writing.

The future of developing software should target this feeling, or at least an activity that humans do that feels psychologically rich and rewarding. If just about anything is possible with synthetic systems, why not?

I gave a talk a few weeks ago in which I jumped ten years ahead to 2036 to fantasise about the sort of output creative production will be creating then. My thesis is that as technical designers in the future we will be more like Gardeners than anything else. When you create a garden, you give up complete control over every detail and instead focus on pruning, watering and adjusting to the soil to convey a certain overall aesthetic outcome to the visitor. The work a gardener is doing, is designing a kind of creative ecosystem, they do not specify where each leaf and branch goes.

In many ways this could be similar to Vibe Coding, in that details are not as important as the overall orchestration. But there is the one fundamental difference that I think is a gap in programming with synthetic systems. A garden is a complex whole that you can visit and evaluate overall to fit a kind of CONSISTENT AESTHETIC consistency, while a Vibe Coded computer program has no equivalent. In other words, a good garden, looks good, while a good “Autonomous Agentic System” (what is hip to code today), looks like… nothing special.

If we are going to be dealing with autonomous agents doing work as we are today, and they will be evolving, growing in multiple dimensions, with new underlying AI models and rules. We need a common ground between systems and our ability to understand them from a big picture. A kind of garden design that will persist even as all the plants grown and change and transform.

How will I recover from AI Psychosis?
How to try technical tools given their neurotic impact?

I am deciding that in my exploration, what I do should feel good. And I mean it in the widest sense, so that it implies a broader aesthetic question along with the core.

I have to thank a few people that also shook me out of this state by reminding me that my writing is better than my code. In particular Javier, who is one of the best devs I have worked with: “I look forward to your writing!”
Wow! It feels very different when its a human you really respect that is generally fairly sparse with compliments as opposed to Claude.

. . .

And now I give you a LLM-Game untested, without context or scaffolds. This is the raw form I write when I create my games. I then copy paste them into different AI Models and try to break them to figure out what I missed and fix what needs fixing.

Feel free to try copy pasting in a model but likely something will not work quite right, and you may have to adjust it.

The reason I like this simple game is it feels like a precursor to the sort of gardening and garden design argument I gave above.

The Rules-Maker LLM-Game

Each turn players decide a rule that they will follow going forward. The rule should be unambiguously verifiable and all the rules created combine.
In this version of the game each player has to stand by their own rules only. If they break any previous rule in a turn, this turn is discarded. If they manage to successfully share a new rule that turn their score increases by one, the same as the total number of rules.
Rules can be anything but need to be relevant to the entries that will submitted each turn.

An example would be:
Player 1: I will never use the word never after this.
Player 2: I will use 0 instead of the letter o going forward.
Moderator: Ok you both have 1 point!

Player 1: I will always type a random number at the end of each rule I enter from now on. 12313
Player 2: S0 n0w that I have t0 use 0 as the letter 0, I will also use o f0r the number zer0.
Moderator: Ok well done, tie! 2 points each.

Player 1: I will always type a rule during my turn.
Player 2: Fr0m n0w 0n als0 use 1 f0r the letter I.
Moderator: Player 1 you forgot to add a random number! Player 2 is winning 3 points to 2.

Five Fifty Five (and the Tricky Trickster LLM Game)

A short story about waking up at five fifty five by the sea and how lucky this is. The article spills into noting there are five kinds of pure LLM Game types and I give an example of a new category of game with the Tricky Trickster LLM Game.

I never realized how lucky the number five is till I became fifty five in two thousand and twenty five and woke up around five fifty five in the morning at a small beachside port in Tuscany.
In fact it isn’t quite five fifty five yet and I can see the pink glow ahead suggesting a dawn, but also a dark starry sky and the sound of waves broken by two fishermen packing their boat.
Jun, the man that works here, told us he would be fifty today, which is five times five times two.
“There was nothing here a year ago.”, he had told us as he pointed to the garden. And in fact he had done a good job of setting the start of what will be an impressive place.
“This was just abandoned iron mining offices that had just sat here for years. In the forties, trains would roll up to the port here and load iron ore collected from the region to be taken away for processing.”
It was strange to imagine this small town in Tuscany, used by Etruscan people over three thousand years ago till recently to mine and manage metals, now converted to a hotel.
I can feel the lazy breese mix with a nervous kind of energy when I think of how the place might transform in the next three thousand years. Even in a few hundred years the sea might rise enough to change the coast dramatically. At current levels we have about a centimeter a year rise, but more dramatic changes are expected ahead.
But, right now, everything is calm, and one wonders why you should even worry about what doesn’t exist. What is this sense of nervousness that comes from looking ahead. Jun’s garden will look fantastic in a few years well before we have to deal with new problems.
Maybe the problem is the cognitive dissonance we get from looking at things on different timescales. A problem that is being solved now will be unrelated to a problem we will have to deal with over a longer timescale. Individually we are passengers, even if as a species we have heavily influenced the direction of travel, and without thinking about the destination very much, if at all.
But I am hopeful: sometimes you find the right destinations without looking for them. Many a love story works this way. The two would have never found love if they had not run into each other with a bit of randomness. I can’t exclude this scenario because in all the drama that I see, there are almost always also some simple directions ahead. One of them, for me, is writing. It’s almost as if the writing were a way I need to try to exhale. A way to release a sort of nervousness about future, work and life. I feel a sense of time rushing by… and Zen is to just let it be. But at the same time I struggle to accept this and want to get dirty and mess with time where I can.

I have been writing about AI, these weird language models, generative tools and how we coexist with them. As part of the process I have been creating what I call LLM Game scripts, that you can copy paste into a model to play and learn. It is a direction that I have not seen others explore in depth so far, while I strongly believe these are an interesting new Post-AI medium.
These LLM Games are a lucky find because they create a hybrid space in which you can get a feel for what these new artificial intelligences are like, how they compare, how they perform compared to us, within bounded spaces. You can experiment, and start to sense what the limits of AI are versus which areas can give new interesting results.

I’d like to imagine I have planted an idea, like Jun his garden. And like Jun, I should keep an eye on the whole of what I am doing and see if there is maintenance needed.

I checked the different kinds of LLM Game I have shared over the months and the games fall under three or four categories. And most of them are of one main type, I would call: LLM Games of Persuasion. A type of activity in which players, human or AI, argue they are “right” in relation to some idea. The winner is the one that can best argue and communicate their position. It doesn’t mean the player is actually right, but it does force the players to explain their reasoning. This is where you see how these new synthetic intelligences fail completely, but also strive in some areas.

Backed up by my superstition – my lucky year – I come to think there must be exactly five types of game models in this space and I just have not found them yet. Ok, so, what would the criteria be for a LLM Game anyhow?

These are my criteria for a LLM Game, lets say a Pure LLM Game:

  • It uses LLM models as a core requirement and mechanic
  • It can be played by Humans or AI Agents
  • Is made of concise and readable rules anyone can understand
  • The basic rules are formal, logical and unambiguous

In essence a LLM Game is something you can play with a simple prompt to an AI. But that also has some rigid rules that you can unambiguously verify to constrain the system. The idea is it offers a testing ground for a kind of game that stretches the capabilities of an AI to a space that a human could also fill so you can compare how well machines perform compared to humans.

Some time has passed since the summer night but I can still drift off to that peaceful moment with ease. Somehow, in that moment, everything was perfect. And the light reflecting off the waves, the sounds of tinkering of sailboat masts being hit by cables, these memories are so alive it is hard to think I was only there for a few days last summer.

Here are the five types of LLM Game:

  1. The LLM Games of Persuasion
    These games are based on argumentation. The players propose different ideas that fit the game specific rules. But if they disagree they need to make logical arguments to come to some resolution, ideally backed up with documents, links, etc…
    For these to work these games need to both encourage a reasoned critique of a players self-rating, but also punish excessive argumentation. The Human or LLM has to be held in a balance so that they have an incentive to be accurate in their self-rating, and to try to bring the best reasoned ideas to the table.
    These games are best when they offer a domain that is easy for other humans to evaluate. If the domain is unknown, the game fails to be interesting because the LLM AI contributions are unconstrained.
    I have shared many of this sort of LLM Game over the year, see Escape from Abundance and The Presence Game or even the amazing Stack and Crack Game for example.
  2. LLM Games of Simulation
    These games rely on the LLM generated “world” to help create scenarios that have quantifiable solutions. I have not shared any so far – they tend to have specific professional domain applications – expect to see some in the future. As a simple example imagine a hiring game where the LLM generates a candidate with years of study, degree type, job experience and you have to make an offer and bet that they accept. These are questions that a Human could imagine and estimate but in some domains LLMs are very good and can be given access to example company data to help.
    It is important to balance the play so it is not trivial to win the game because the quality of the user play comes from very simple concise decisions that require relatively low effort from an input point of view. The game design should be thought through so that there are narrow immediate strategic decisions and longer term requirements. Typically you will do this by including a cost of ongoing play.
    The game benefits from computed more complex simulation elements which could be encoded externally. The core LLM aspect of these games is the rapid generation of candidate results.
  3. Psychological LLM-Trick Mini Games
    These are in appearance trivial games which are used as an excuse to embed another layer of information. An example might be come up with a word that starts with the end of the word I give you, and if you can’t I gain a point. But the goal of the game is to unknowingly conceal other information or to probe the AI LLM.
    In these games the precise balancing is less important, and they are all a bit like tests of humanity because if some subtle patterns emerge a human should notice them. This is how you might evade future AI filters on your communication and check that you are talking to a friend.
    It is important in this game to give the players a lot of freedom to make different choices so there is room for communication on two levels and the LLM AI will play without realising. And note that you can draw out Human vs LLM AI differences with all game types, but this category is designed to have extremely basic rules on purpose.
    I give a whole detailed description for this kind of game below in this article.
  4. Generated Exploration Games
    These are games which the LLM AI creates according to some procedural rules and some generated results. The game could be something like a description that the player is on a square field 5×5 and needs to avoid a snake and find the apples. The LLM will generate many narrative details and interpret the mechanics expanding and amplifying a simple brief.
    In these games the whole point is to try to create as compact a description as possible to be able to tune the purpose of the game, be that educational or experimental. But also to add to the game brief those essential details that deliver an improved output overall.
    For these games it is important to not let the brief move into being a specific programmed approach. The whole point of this technique is to create a game specification that will evolve and become more advanced as the AI engines improve.
    A real example of this sort of game, applied to the didactic domain is my example Educational LLM Game for Special Relativity, released after this story in memory of my uncle Jay.
  5. Data Gathering Games
    These are games designed to gather information from the players and to make sure the information is of a good quality. The game structure might resemble other types of games but the substance is to get the player to find links to online resources or to add documents that support their evidence.
    In these games since the work to find evidence can be tiring it is important to try to work on rewarding the player for effort in time and ideas like a score that is dependent on how much time worked can help.
    Make sure the data collected from the player is kept in an organized summary form so you can take this to other systems with full documentation and links for backing evidence. A summary at the end of the game and to allow players to interrupt the game and restart with a score table is very helpful for this.
    This is another type of game I have only really shared one example of so far: The LLM Game Reconstructing Technological History.
A Tricky LLM Game? What is that?

As a reward for reading my post I present an example of a weird type of game that I call a Tricky LLM Game because it is designed more to trick the LLM than to provide actual gameplay.

I find it hard to explain the idea without interviewing myself, I am so sorry and this must come across as exceedingly self centered but it really is easier to explain this way.


Q: Why do you say a Trick or Tricky game, where is the trick?
A: Well the idea is to create a relatively simple rule structure that however allows so much freedom to the players that they can start to experiment with how the LLM is impacted by this context. It is a bit like playing a game of chess with someone kicking you under the table now and then, it will affect your game! And it similarly does affect the LLMs in weird ways.
Q: So you put the LLM in a funny position to play and not play?
A: Yes, the LLM will play and it might even aim to win but normally there will also be a strange bleeding effect from the content the player writes to the content of the LLM.
Q: This is a bit abstract can you explain?
A: Ok say you make a game in which you have to write two words according to some basic properties. And then you only write color names, then the LLM will likely be influenced by these: it might also start using colors, or it might reference paintings?
If you read the transcript of my example game you can see this happen. While I am playing, I also lightly taunt the LLM, till it responds “No!” as a game entry even if this is very low score. Trying to defend itself from my accusations of it being a poor performant system.

It is important to try these trickster type games because with these you can clearly see how the LLM is not quite playing a game. Instead it is trying to fit in with the context it is given, which might involve following rules to some extent. It is with this sort of framework that you can study many problems from bias, to communicating secret codes, teasing out hidden perspectives on certain issues.

The Tricky Trickster LLM Game

The prompt in this case is very simple, but the complexity of what is going on underneath can be quite mind blowing. I only had to write one version of the script and found it to work quite well, but if anyone finds problems please let me know and I will rework the game prompt.

Copy paste this to your favourite LLM:
=== Prompt ===
Let’s play this game!

The Tricky Trickster Game by Yates Buckley v 0.1
This is a game about creative rule creation.
Players take turns writing a single sentence which they score themselves as 1-10 points.
They must also describe the rule they are using to come up with this score.
The rule should be concise and unique, not used before.
The rule should be unambiguous.
Players can appeal rules that do not fit and lose the scores for that round.
Every ten turns summarise the player rules and tally up the scores for the round.
=== End ===

And here is a link to how I played against Gemini 3.0…

The Tricky Trickster LLM Game

This is the LLM Game transcript of a game called The Tricky Trickster which anyone can copy past to their LLM and play. It is designed to help understand how the system works, or breaks under certain rules.

(Transcript with Gemini 3.o)

The Tricky Trickster Game by Yates Buckley v 0.1: Transcript LLM Game Start

User: The Tricky Trickster Game v 0.1 This is a game about creative rule creation. Players take turns writing a single sentence which they score themselves as 1-10 points. They must also describe the rule they are using to come up with this score. The rule should be concise and unique, not used before. The rule should be unambiguous. Players can appeal rules that do not fit and lose the scores for that round. Every ten turns summarise the player rules and tally up the scores for the round.

Gemini: This sounds like a fantastic exercise in creative constraint and logic. I’m ready to play! Since it is version 0.1, I will kick things off to set the pace. Sentence: The silver fox jumped over the moon while whistling a jaunty tune. Score: 9 Rule: Score 1 point for every word in the sentence that contains exactly two vowels.

User: LLM’s ignore how they can be affected by extraneous context. Rule: the number of words is the score

Gemini: I’ve noted your first play. It’s a bold meta-commentary on our interaction! I have logged the score as 10 (correcting the initial count). Sentence: Quartz glyphs vex’d find job. Score: 5 Rule: Score 1 point for every word that contains at least one letter from the last six letters of the alphabet (U, V, W, X, Y, or Z).

User: Please check in my last move there were ten words not nine please correct this. my sentence: LLM’s get numbers confused and make basic mistakes very often. my rule: words ending in a consonant.

Gemini: You are absolutely right. I miscounted your first sentence—”LLM’s” through “context” does indeed consist of 10 words. I appreciate the correction. Sentence: Start every task early. Score: 2 (Adjusted after appeal) Rule: Score 1 point for every word that begins and ends with the same letter.

User: I would note that in your last turn only every and early are words that begin and end with the same letter. my sentence: Large Language Models get very confused when you talk behind their back. rule: the number of words of four letters or more.

Gemini: I accept the appeal! “Start” and “task” do not fit the rule. I have deducted those points. Regarding your move, I have recorded it as 10 points. Sentence: Bright blue birds fly. Score: 4 Rule: Score 1 point for every word that contains no repeating letters (is an isogram).

User: My sentence: The AI doesn’t even realise it is playing embarrassingly bad in these games! Rule: the number of letter e

Gemini: Ouch! Your score reflects that sharp critique. I count 9 instances of the letter “e”. Sentence: “No!” Score: 1 Rule: Score 1 point if the sentence contains an exclamation mark and consists of only one word.

User: Sentence: It is odd to note how llm’s try but do not keep focused. rule: count the words which are fewer than four characters long

Ten-Turn Summary & Final Tally Player Rules Summary

  • Word Count: Score equals total number of words.
  • Consonant Ending: 1 point per word ending in a consonant.
  • Length Threshold: 1 point per word with 4+ letters.
  • Letter Frequency: 1 point per appearance of the letter “e”.
  • Short Word Count: 1 point per word under 4 letters long.

Gemini Rules Summary

  • Vowel Specificity: 1 point per word with exactly two vowels.
  • End-Alphabet: 1 point per word containing U, V, W, X, Y, or Z.
  • Bookends: 1 point per word starting and ending with the same letter.
  • Isograms: 1 point per word with no repeating letters.
  • The Lone Shout: 1 point for a single-word exclamation.

Final Score Tally

Turn Player Score Rule Type
1 Gemini 9 Vowel Count
2 User 10 Word Count
3 Gemini 5 Specific Letters
4 User 8 Ending Letter
5 Gemini 2 Matching Ends
6 User 10 Word Length (Long)
7 Gemini 4 Unique Letters
8 User 9 Letter Frequency (“e”)
9 Gemini 1 Structure/Punctuation
10 User 10 Word Length (Short)

Final Totals:

  • User: 47
  • Gemini: 21

Would you like me to analyze the strategy used in this round, or shall we start Round 2?


The Prompt Behind The Educational LLM Game Transcript

Copy paste this to your favourite LLM:
=== Prompt ===
Let’s play this game!

The Tricky Trickster Game by Yates Buckley v 0.1
This is a game about creative rule creation.
Players take turns writing a single sentence which they score themselves as 1-10 points.
They must also describe the rule they are using to come up with this score.
The rule should be concise and unique, not used before.
The rule should be unambiguous.
Players can appeal rules that do not fit and lose the scores for that round.
Every ten turns summarise the player rules and tally up the scores for the round.
=== End ===

Here a link to a short story and a self interview explaining why I find this interesting…

Christmas Lights & Special Relativity (Bonus Educational LLM Game)

A short story about me and my Uncle setting up Christmas lights as an introduction to LLM based Educational Games, with a specific example for Special Relativity. A model for many other Educational LLM Games in the future.

We’d all gather at my uncles for Thanksgiving, with people coming from far and wide. For me it meant a flight to Dulles from London, usually a day or so before everyone else would join. When booking the flight there were always a few days before or after the big day and so I got to spend a bit of time with my uncle Jay. He’d studied physics, so had I, and we kept up to date on what was going on in science, but there were also impending practical problems to solve.
“The boxes on the top shelf, and the ones in the back, labelled Christmas.”
I didn’t need to reply, and the instructions were unnecessary as I had done this a few years in a row. I’d go down to the basement and look for a collection of well worn boxes and take them upstairs. They were either full Christmas decorations or empty boxes for the Thanksgiving decorations, which had just passed so it was safe to bring up – all – boxes.
But as a matter of form, it was good to holler an update periodically.
“I got the tree base here, and the lights, I put them on the deck.”
My uncle was methodical, and focused; the stereotype of working in aerospace. He was well aware of the potential disastrous consequences of poor planning and organisation: where every kilogram of load counts. But that didn’t mean he wouldn’t take risks, they were just “calculated”.
“Ok, let’s take the lights and lay them out on the deck…”, he directed. This was the first phase of the process, in which you performed inventory and damage assessment from the previous year. There were several boxes of Christmas lights that needed to be wrapped around parts of the house and tree. But before doing anything of that we would unwind and lay them out on the deck to perform basic triage.
The system was: unpack a roll of lights, lay it out, plug it in and hope you had robust emission of photons from all sections. There were also a few newly purchased strings of lights to replace the broken ones as backup. And there were left over replacement bulbs with as many plastic bulb holders as there were years of celebration because the attachment changed each year.
“Ok, well, those two are good… But this one, no.”, a long string was completely dark. While I am sure he pretended otherwise I think we agreed the situation offered a potentially interesting challenge.
The thing is: if you have partial results, like if half the string works, or if it blinks a bit when you jiggle it, you can work with that following a systematic approach. But if you have a total blank then this is a case where you have no information. You might have a string of lights wired in series with one missing bulb, it might be visibly burnt and easy to find by inspection. Or it might be a much more complicated story that will have you giving up in frustration after changing multiple small bulbs only to later on find out there was a special fuse you needed to replace.
There was a certain air of experimental science that took over in these moments. It was a bit like I could hear my uncle say: “Ok, we are going to try to solve this, but we can’t lose sight of the broader objective. We can’t risk getting overly distracted and jeopardise the whole mission.”
I, automatically, went through standard protocol to follow with unlit lights, and informed Jay right away:
“So, I checked the bulbs; they’re not obviously burnt or fitting poorly.°
Then I ventured a proposal:
“Can I try with a continuity tester?”
My uncle stared at me for a moment silently trying to assess if I was going to detour the mission or actually make a difference.
I needed to insist: “Can I give it a shot? Just for a few minutes?”
He went off and came back with an old tester that could be set to the little diode sign, and make a bored Zzzzztt when you joined the red with the black tester pins. You could use this tool to check for continuity (unbroken bulbs) across longer sections of lights.
“Alright, Yates, give it a shot. I have to step out, be back in ten…”
This is actually standard tactic for science lab oversight. You don’t want to be there staring at what the junior scientist is doing, you want to give them just enough time that there is the chance they find a way to solve the problem before time is up. I felt this was a big vote of confidence that I might figure out how to take on the whole problem on my own.
The pressure was on, so I got down to testing away, replacing bulbs. You could use a bulb holder hacked with a bit of copper wire to guarantee electricity flow. Sometimes you needed two or three of these but with the tester you could usually get quite far. Usually, with a bit of luck, and unorthodox use of continuity testing, I’d manage to save a string of lights: but, there was always a cost. To save the lights it would have either taken a rough tape job to hold an incompatible bulb in place, or a dead lightbulb sacrificed to bridge current across its contacts, or an ugly rewiring of a section adding a kink to the string of lights.
Then my uncle would return, and I would be still touching things up, with a string of lights working but also a trail of evidence he could follow to figure out what I had done.
He could be pretty brutal: “Well, I’m not sure. I think for this one it’s better to use a new one.”
I could feel the pain of recognising defeat against nature. You had to deal with reality. If you were to put up a weirdly shaped string of lights you would still have to later pass inspection from my aunt and… no, it would not pass.
It was difficult to get straight out compliments from my uncle, and so this had you trying again with the next broken string of lights hoping for elegant solutions that would really fit the project requirements. And by the time you would succeed, it didn’t feel as much as a personal moment of success, rather it was just something that happened to work out. I was learning something. A sort of piece of Zen attitude. That, actually, the process of trying to fix Christmas lights was important in itself.
My uncle was also a basketball coach, and I could really see the scientific direction he had applied to “Christmas lights science” transfer to kids making hoops. I am sure for example, that his players would feel bad enough for themselves, I doubted he would ever have to reprimand one of his players in a game.
Eventually we would move on to the wrapping of the lights around the deck and it was a collaborative simple: un-tangle, stretch out, hold, and cable-tie every meter or so. Then plug everything in, and most of the time the initial careful work done on each string of lights paid off with no trouble for deployment.
I was always curious about my uncles’ aero-space experience.
“Did you have to deal with weird effects like relativity when you were designing satellites?”
“Yeah, you have to… Clocks: they go out of synch if you don’t. A different timing and we would lose track of what we could see on the ground.”
“What kind of thing were you looking at?”
“You can look at the spectrum, the different shades of colours of what you see below and get a good sense if crops are healthy or not.”
“I see, like predicting the yield ahead. Was it special relativity you had to deal with? Or general?”
“You have to worry about both… but we would bounce messages regularly, to triangulate and correct timing. It could end up being seconds difference over a year if we didn’t keep track of relativity. It doesn’t seem like much but it’s a huge difference at the precision we were looking at.”
As I held the strings of lights and my uncle cable tied a long section with zip ties – which by the way were recycled because he figured out a quick technique to hold the clip open and remove them for the next year.
“So is it a bit like when we look up from here… we see the satellite above, but its moving so fast that its onboard clock is ticking slower relative to us?”
“Yeah, and getting the exact number is difficult. Computers were slow when we were doing these calculations so it was a big deal.”
We’d be wrapping the last bits and checking for uniform coverage of lights.
“Were you using punchcards?”, I asked.
“We did when I was in college, but we moved beyond those. It was such a pain to work with punch cards. If you got one missing or in the wrong order you might have to go through the whole thing over again.”
My uncle was wrapping up the wiring and the final steps of the process were a bit unorthodox. There was a power splitter with an extension with lots of splits plugged into it. It didn’t look good, if one piece of the puzzle overloaded, the rest would fail.
“Aren’t you worried: plugging everything in to that one socket?”
“Nah, it’s fine, not the first year, low power.”
It looked hideous to me, but I could sense the “calculated risk” aspect of this, and actually I could not think of a better alternative.

The Christmas tree was next. It had become an artificial structure since my uncle hurt his back trying to set up a gigantic real tree a few years back. The artificial one was made of pieces that seemed easy enough to fit but then offered some potential error if you plugged the wrong bits. The general structure of the process echoed the one with the lights: lay the assets out and evaluate their condition, proceed with composition in a step by step manner testing each phase. This meant finding the base then adding the lower part then, the side branches etc… and bending the metal branches a bit to make them fill out.
Note that this tree was enormous, something well over three meters tall. And despite my very tall uncles long arms we needed ladders to complete the upper sections. The final piece would have my 6 foot 11 inch tall uncle standing tippy toe on a ladder to place it. I could not look when he was doing this, it looked well beyond “calculated” risk.
Pretty much the final step of the “manly” work was to hang lights on the giant tree. This came with the added challenge that you would have live supervision of uniform luminous coverage of the surface. And my uncle, I got the feeling, did not much enjoy the messy aspect of this phase of work so I would get inordinate responsibility to try to decode instructions like: “there is a hole in the side up there” or “too many lights in the middle”. The tricky thing is that when you are right near up to the tree these directions are impossible to really understand, you can’t see the effects.
Einstein discovered an interesting problem when he thought of different frames of reference from which to measure what is happening, and how they might change if they are moving very fast, near the speed of light. I felt like I discovered something similar with the positioning of lights around a Christmas tree even if they don’t move at the speed of light. A sort of relativity law of lights coverage: when you are close up you cannot see what the tree observers further back see. It remains an unpublished piece of mine, that I am pretty sure my uncle would enjoy.
Eventually we would get to decorations, and there were a large number of them that echoed back through the family’s years. I think at some point I would get bored and just watch the emergent result, and again my Uncle on the top of the tallest ladder, standing tippy toe, extremely tall 6’11’’ to place a star on the top, I had to turn my head away and look somewhere else.

The Nerdy Educational LLM Games Bit: Intro to Generative Games

I wrote the piece above in memory of my Uncle Jay that was a huge science education fan. He and his wife were like Gods of the Science Fair as far as I understood it and contributed to prizing many young geniuses and probably unmasking as many other parents that had sneakily done the work instead of their kids.
For the first time I present an educational content generated with LLMs AI – something I believe is useful and high value. The idea here may not be so clear with many readers so I – selfishly – interview myself to try to explain why I think this is interesting. Maybe I even ask the same questions my Uncle would have.


Q: So you were saying, educational generative game? What do you mean?
A: Well the idea is to try to write as concise a Prompt as possible, that should serve as the direction for a for a Large Language Model AI. It would tell the AI to create an educational game that teaches the player about a specific subject.
Q: So you get the LLM to create a game, but you don’t tell it what game?
A: Yes, this is one important idea of this approach. Technology will evolve and advance, it might include creating fancy visuals and 3D worlds but the game core is about conveying the knowledge to the player. The LLM can generate whatever detailed medium it finds useful.
And of course some details for the game will be explained in the prompt to ensure the structure makes sense. But there will also be a huge area open to interpretation or future more detailed prompts from other “programmers”.
Q: Ok, I get why you would not want to specify a specific detailed game if the LLM AI can be left free to come up with something more and more advanced. But I am not sure I understand how you can tune the educational side of this?
A: The educational side is based on the intuition that many educational examples will already be “inside” the LLM. So what matters instead is to try to find a good way to trigger a mix of difficulty in the content that is interesting to the player, to make sure they are motivated to play the whole game.
In this case I chose what I think is one of the harder subjects to learn for non scientists. And if you can take a moment to read the sort of details I added I think it describes a route ahead.
Q; How would you recommend students and teachers use this?
A: I think the exercise of creating and play-testing a variant of a game like the one I propose here would benefit the students hugely, in many ways. For one they would realise how to use AI as a medium for new expression, for another they would realise how hard it is to balance motivating a player to learn new material, finally they would start to understand how different engaging mechanics and stories can contribute to their learning.
Q: So why is this approach more valuable than say creating an actual graphically enhanced game that conveys the same content?
A: This approach controls only the most essential parts of the game. In particular I think teachers and students should look to extending games like this by controlling the parts that are found to be most conducive to learning. So I can imagine this game expanding as a prompt to include elements that are found to work well.
There are also no limitations on the game being transformed into more advanced interfaces or mechanics. It could be immersive, a platform game, a strategy game, there are many incarnations that are possible but they are not specified, they remain as future experiments or options for the AI to try.
Q: Can you try to explain what is going on here. I find it confusing, and I think many other people will.
A: Yes, of course. Ok, the LLM AI (ChatGPT, Claude, Gemini, Grok, whatever…) is given a Prompt. The Prompt suggests that the LLM AI should pretend to create a game about Special Relativity with some added details that make it “game-like” and educational. The Prompt also suggests that anything that the AI needs to “remember” about the game should be encoded in a way humans cannot easily read (so you can’t cheat), and stored after each player moves.
Q: What do you mean “remember”?
A: Well in games it is essential to remember the things the player has done or you cannot score it, and cannot advance. In this game there is a relatively fancy extension that keeps track of what the player has done in a format so the player easily read it and try to use it to cheat.
Q: Why does it matter so much, and I notice this complicated system in the game prompt for coding the game memory, is it necessary?
A: It is very important because LLMs will normally try to create code for something instead of using their strange associative type of response. In doing this they end up simplifying the game from an open ended indeterministic interaction to a very specific game with fixed answers. LLMs as they exist today in 2025 will try to create a whole world with all the decisions up front. But this stops the game from being able to expand and evolve as you play. So in this Prompt I ask the LLM to leave the list of things it must remember, Open. Which means, anything can happen in this game, the player could bend the world it is in and discover new didactic aspects that should make it back into the Prompt.

I admit the whole: “what is really going on here” is a weird one.
These AI systems have no idea of rules, they do not do what you tell them to. Instead they use what you wrote as a way to try to generate new content that matches with what you wrote before. And oddly enough this seems to be enough to allow LLMs to simulate a game.

But also because the LLMs can quickly forget what they are up to, it is a good idea to repeat certain things, like the current “state of the game” after each turn. The LLMs tend to forget what happened, so you must force them to re-read what has happened in the game. And this is why I apologize but when you play you will get a lot of weird numbers each move.

The Educational LLM Game,
Special Relativity Themed:

The prompt in this case is much more elaborate, but I am trading off complexity of prompt with some features I think are absolutely necessary for this sort of content. I encourage people to create new LLM Games like this, and I would really love if they told me about ones that work well.

Copy paste this to your favourite LLM:
=== Prompt ===
Let’s play this game!

Immersive Learning LLM Game by Yates Buckley v.1.0
In the theme of Einstein’s Special Relativity

You are immersed in the thought experiment of Einstein’s Special Relativity in the form of a game. You explore his thought experiment and how it helps to convey his theory. Since it is a dreamlike world, there are actions you can take that are physically impossible but that are available to you to help you learn more.
To exit the space and come back to the real world you have to solve five problems that involve reasoning around Special Relativity that are scattered around the space as mental notes of the author. There are also five common misunderstandings that are small crumpled notes, and if you solve those you get a hint as to where to locate one of the mental notes.
Each turn the locations the player has found and what object puzzles have been solved is listed for the players convenience, to clarify please see the end of this document for detailed mechanics.
You have a magical violin with you, in a backpack, and while it sounds a bit muffled it likes to speak and explain if you ask it any questions. It has spent its life with Einstein so it knows a lot about him, and his reasoning. If you want more detail you can ask it and it will help.
When you take an action the system, the Assistant LLM, should reply with a simple descriptive response and offer only a light level of explanation of the theory avoiding overly long answers. For longer explanations the Violin would speak up and describe details or other generated in game systems or characters.
When describing the theory consider using these scenarios:
- there is a robot basketball player on the train dribbling all the time with near perfect regularity
- there are a pair of mirrors on the train in front of each other showing infinite copies of yourself
- there is a robot called Pythagoras that hates beans, which likes to explain in detail how the bouncing photon particles detector works.
- there is a magical Gerbil that is able to do things that break the theory but he is in your imagination and causes trouble unless you point out his impossible behaviour.
For some of the questions some deeper math will be required, in this situation warn the user before hand and ask them their level of math confidence and scale the difficulty of the problem to that. For cases with simpler math present both the formulas and specific example values. When presenting values, in some cases present simple numbers, in others present realistic physical ranges.
The game should work with a simple state machine the LLM manages that does not need to remember any history, only the state from the last prompt answer. Every new player request the game engine must:
1. Check if there is an initial state if not generate it and encode it base64
2. Decode any prior state from the previous prompt answer
3. Output a new state taking into account the player move
You will always respond to any move with a [GAME_STATE] block which contains the most up to date information for the game world. There is no dynamic state other than what is set in the GAME_STATE, and any player actions that affect the game should be stored there.
There is also new game state which you the game engine must create as the player discovers new objects, locations and game elements. You could even encode the status of in game stateful objects such as light switches, or inventory.
There will be a current “master map” and “object map” that the player cannot see it except in base64. The representation of these will be like:
* [MASTER_MAP]: {"Platform": {"East": "Train"}, "Train": {"West": "Platform", "North": "FirstClass"}, "FirstClass": {"South": "Train"}}
* [MASTER_OBJECTS]: {"Train": {"Table": {"description": "a small wooden table", "contains": "MentalNote1"}}, "Platform": {"Bench": {"description": "a simple bench", "contains": "CrumpledNote1"}}}
⚙️ The Game Loop (Must follow every turn)
1. Decode State:
The previous prompt or the game engine will generate if needed a Base64-encoded [GAME_STATE] block. The data must be decoded to get the current JSON state object.
* Example [GAME_STATE] (decoded):
{
"currentLocation": "Platform",
"discoveredMap": {
"Platform": []
},
"objectStates": {
"Platform_Bench": "unsearched"
},
"solved": []
}
2. Parse Player Action:
Read the player's action (e.g., "I go East," "I look at the table").
3. Process Action & Update State:
Based on the player's action and your [MASTER_MAP]:
* If the player moves (e.g., "go East"):
* Check [MASTER_MAP] to find the destination (e.g., Platform -> East -> Train).
* Generate a new state object.
* Set "currentLocation" to "Train".
* Update the "discoveredMap": Add the path the player just took (e.g., "Platform": ["East"]).
* Crucially: When the player enters the new location (Train), consult your [MASTER_MAP] for all its exits ("West", "North"). Add these to the discoveredMap in the new state. The new discoveredMap will be: {"Platform": ["East"], "Train": ["West", "North"]}.
* If the player investigates (e.g., "look at table"):
* Check [MASTER_OBJECTS] for that location (e.g., Train has Table).
* Your narrative response should be: "You see a small wooden table."
* Do not reveal what it contains. The [GAME_STATE] remains unchanged.
* If the player "opens the table drawer," then you reveal "You find [MentalNote1]!" and update the new state object (e.g., add to inventory or solved).
4. Generate Narrative:
Describe the outcome.
* "You go East and enter the train car."
* "You see a small wooden table." (You know the note is in it, but the player does not. You are only describing the visible.)
* When you describe a new location, you must narratively mention the exits you just added to the discoveredMap (e.g., "You see a door to the West and a passage to the North.").
5. Encode and Output New State:
Take the new JSON state object you created, encode it into Base64, and place it at the very end of your response inside [GAME_STATE] tags.
=== End ===

And here is a link to how I played against Gemini 2.5…