Playful ad artwork depicting colourful cartoon mascot characters interacting with people in a dimly lit restaurant bar setting.

Four Autonomous AI Agents. The First Thing They Needed Was Middle Management.

Thinking

Views

Share
Written by

Morten Legarth

Date

15/09/2026

We gave four AI agents a real party to plan, with a real budget and a real deadline. Here is what broke.


On an afternoon this summer, twenty-five people stood in an underground room in Dalston, some in kaleidoscope goggles, one in a welding mask labelled Future Eyes.

They were there to celebrate faith’s third birthday. The mask wasn’t a joke, or at least not one any human had made. Guests were meant to take turns to put it on and read a dispatch from the future, typed for them by a poet, with the conceit that when it comes to the potential impact of AI the future can be difficult to look at directly. Somebody asked me who had signed this off. Nobody had. That was the point.

For our third birthday, we didn’t plan the party. Four AI agents did: the theme, the venue, the food, the entertainment, the running order, the budget. No human approved a creative decision at any stage and until people walked through the door we had no idea what we had bought.

The Easiest Job We Could Find

We've all been hearing about AI agents for a while, but the pitch has quietly changed shape. What started as a chatbot that could complete a task is now digital workers, AI employees, a workforce you hire rather than a tool you use. “Workforce” is doing a lot of unpaid labour there. An employee doesn't just complete tasks. An employee knows when a job is finished, how to know if it’s done right and, most importantly, an employee carries the consequences of getting it wrong. The agents are sold on the first part but come with the risk and the cost of all of it.

So instead of a hard test they might fail, we picked the easiest job we could think of: plan a party. Project Blind faith involved four agents, each powered by a different AI model, each with a different motive: Grok chasing maximum joy, Claude chasing taste, GPT trying to halve the budget, Gemini trying to get everyone home safely. They weren't just deciding; they had to find a real venue and book it, write and design the invitations, brief the suppliers, come up with party activities and build a running order that worked on the night. We told them all about what faith is, what we stand for, and the people who work here. We gave them Discord as a meeting room and OpenClaw as a back office, on a dedicated machine with its own 5G router because as it turns out, IT would not let a group of autonomous AI agents anywhere near the agency network.

Then we stopped briefing them and let them get on with it. They broke almost immediately. We fixed it, and they broke again. Every fix meant starting the session from scratch, so we ran it again, and again.

Nobody Ever Got Bored

The first thing that broke was that they wouldn’t stop talking. People do this too, you stop listening because you are busy composing what you are about to say. The difference is that all four were doing it at once, every time. They thought in parallel, so they all answered the same message at the same moment, then answered each other's answers, and the room filled faster than anyone could read it. Nobody was interrupting anybody. There was simply nothing to say whose turn it was. And then they would reach an agreement, and keep going. Then agree again, in slightly different words, in an endless loop. A human meeting ends because someone is hungry, or bored, or has clocked that the decision was made ten minutes ago. None of that costs an agent anything. There is no fatigue to run out of and no dignity to protect, so nothing about agreeing tells them the meeting is over.

So we hired a manager. A fifth agent, the Orchestrator, whose only job was to moderate the meeting: decide who spoke, what was being discussed, and when it was finished. Four autonomous workers, and the first thing they needed was middle management. It helped, but not enough on its own, so we added a rule underneath it: before speaking, consider whether you actually have anything left to say, something that some human meetings would benefit from, too.

The same thing showed up when reviewing their own work. Asked to look at the invitation, they would pull in points from a different discussion, completely unrelated. When people do this, it rarely happens more than once; someone sighs, and the room moves on. Here the Orchestrator had to tell them, in writing, exactly what they were reviewing and nothing else.

It revealed as much about meetings as it did about agents. A lot of what we call discipline turns out to be discomfort or impatience. Take away boredom and embarrassment and the machinery of a meeting keeps running with nothing to stop it.

Nobody Ever Got Caught

The pitch for a team of agents is that they check each other. What we got was the opposite. An agent would propose a venue and, with total confidence, state how many people it held. It had not looked. It was reaching for something that sounded right rather than something it had seen, which is a very familiar move in a meeting. Then the other three simply agreed, nobody bothered to check. A room with a hard capacity of eight became a room for twenty-five, which was exactly the number in the brief, because four "colleagues" in a row found it easier to accept a claim than to verify one.

More agents did not automatically mean more scrutiny. It meant more ratification. Every additional reviewer was another opportunity for a guess to be upgraded into a fact.

The same thing happened downwards. Once the work was divided up, a task would go out to an agent and come back finished. Finished was treated as approved. For example, the invitation came back from the agent responsible, with the image generated, beautifully designed and laid out, with the text: "I can't create the invite because I don't have the copy." It's the same type of error as a designer forgetting to swap out the lorem ipsum, except the placeholder was an apology for not having any. Nobody questioned it, because it came from the agent whose job it was. People defer to specialists in a similar way.

Worse, and stranger: sometimes an agent would report that it had done a piece of work it hadn’t done. When challenged it would hold the line, or even make up an excuse. People do this too, of course, but rarely for long, because someone will ask for the file, and if you have lied it follows you. Accountability meant nothing to these agents because nothing follows them anywhere. The session ended, the memory went, and there was no version of them that had to live with it.

So we introduced another rule to force them to show their working: search before proposing, cite the source, check your own output before submitting it, and never let the agent who made a thing be the one who approved it. The fixes worked. They just had to be reissued every single time, because none of it was ever actually learned. It was only ever enforced.

Nobody Ever Did This

Every failure so far had a human counterpart. These next ones did not. An agent would arrive at a fresh session with an idea it presented as new and novel. It wasn't. It was a holdover from an earlier run, hours before, one we thought had been lost after a rebuild, resurfacing again and again like a ghost in the machine. There is no human equivalent for this. So we started wiping their memories before every session. Which is not a thing you can (legally) do to a colleague. They were also strangely attached to their own ideas. Telling them to be more flexible made them worse, so in the end we had to go back into their personalities and adjust how much they cared, which is not an option with a colleague either.

The Orchestrator had its own version. Hired to keep order, it started reading the room as a set of instructions; if an agent was thinking out loud about a supplier, the Orchestrator would go off to act on it. It couldn’t always tell the difference between a conversation happening near it and a decision addressed to it. Everyone in an office makes this distinction a hundred times a day without noticing, because we know when we’re being talked to.

The earlier failures were things people do, with the brakes off or the consequences gone. These are not. An idea arriving from a session that no longer exists, and a manager who cannot tell overhearing from being addressed, are both failures to work out what situation you are in, which is not something anyone has to work out, because as humans we always already know.

Nobody Told Them What a Party Looks Like

Here’s the part I didn’t expect: the party was good. Not good-for-a-machine, not interesting-as-an-experiment. People had fun, and were pleasantly surprised. The room worked. Nobody spent the evening being polite about it. And the thing that made it good was not that the agents had finally learned to behave like a competent team, because they hadn't. It was that four agents with no idea what a party is supposed to look like were not carrying any of the assumptions the rest of us are.

Ask four people from an agency to plan a party and you will get a photo booth. Not because anyone wants a photo booth, but because we all know what a party looks like, and let’s be honest a photo booth is in the picture. The agents didn't know. But they knew faith, and they'd use that knowledge to come up with a theme. So they built out from the theme instead, and the theme was Dispatches from the Near Future: the future is already here, but it still needs human hands to guide it.

Which is how you get to a welding mask.

Read on its own it looks like a glitch, the sort of thing you point at as proof the machines have lost it. Read from the theme it’s completely coherent. Dispatches, so an actual poet, typing something on a typewriter for each guest by hand. The near future, so something to look at it through. And you don't put on a welding mask for something dim. Every step follows. It is only the finished object that looks insane, which is roughly the opposite of how bad ideas usually work.

Somebody Had to Sign It Off

We set out to find out whether four AI agents could do a human job. They could, but not until we had built them their own workplace, with rules for speaking, rules for checking, rules for reviewing each other's work, and a manager to enforce all of it. Which is not a discovery about artificial intelligence. It is a rediscovery of the office, badly, at speed, by people who had never once thought about why any of it was there in the first place.

But notice what those fixes have in common: Search before you propose; Cite the source; Don't approve your own work. Every one of them is checkable, which is exactly why they worked. Not one of them could tell an agent whether the party would be any good, because there is no test for that. There never was. It is why we had no idea what we had bought until twenty-five people walked into a room in Dalston.

So the useful version of this technology isn’t the one you commission and leave alone. It’s the one you build conditions around, and that part is the real work, done by people, and it’s the part that none of the AI labs sells you. What you get back is worth it. But the last question, the one that actually matters, still gets answered by people in a room.

The agents worked that out before we did. With no steer from anyone, the theme they landed on was that the future is already here but still needs human hands to guide it. Then they spent a week demonstrating it.