faith And The Toothbrushing Test
The evolution of AI video, a mundane challenge, and why the future belongs to “World Models”.
We launched our creative studio, faith, in the very earliest days of generative AI video, and we've watched the technology evolve alongside us, from short low-res clips of eerily morphing characters to multi-scene HD clips of Hollywood-level action.
But as the barrier to creation drops, the bar of quality is raised higher; and even with the increased sophistication and seemingly infinite possibilities of today’s synthetic videos there are still limits to be found. And often these limits aren’t the high-end action and emotion required for broadcast-quality production, but something as simple and mundane as brushing your teeth.
Primitive Pioneers (2023)
The first commercial text-to-video model was Runway’s Gen-2, announced in March 2023 and released in June. By today’s standards, its capabilities were modest, to say the least: it could generate four-second clips, with a small amount of (slow) motion. It wasn't perfect—far from it—but people immediately started using it to tell stories. They’d edit the clips together in interesting ways, choosing to either embrace or ignore the (many) flaws.
Already at that early stage, you could see a clear pattern emerging: people with deep experience in traditional filmmaking and video storytelling were getting the best results out of what was otherwise a very basic model.
A few other models, such as Pika and Stable Video, followed in late 2023. These offered modest upgrades, perhaps the most notable of which was the image-to-video capability, which gave creators a much-needed anchor of control. Even so, AI video wasn’t close to being ready for commercial use.
Sora and the Second Wave (2024)
In early 2024, OpenAI announced Sora. It promised to shake up everything we knew about video models, showing off sprawling, realistic minute-long generations.
The announcement came just as we’d started making our short film, finding faith, to celebrate our first birthday. In retrospect, we were incredibly early to the game. Creating the film to the standards we’d set was a battle; one key scene alone involved hacking together motion tracking, 3D modelling, image and video generation, and a lot of painstaking post-production.
We thought that the release of the Sora would mean we'd have to scrap all the work we'd done and start again. But OpenAI chose not to release it until the end of the year.
While we were waiting, we got a second wave of models during the summer and autumn of 2024. This included Luma's Dream Machine, Runway Gen-3, and the popular Chinese models, Hailuo and Kling. These systems brought better detail, improved motion, higher resolution, and longer clips.
It was around this time that I started experimenting with my own personal short films as a way of better understanding the capabilities of these new models. I was always more interested in how these networks handled subtle human motion and emotion than I was in big, explosive action sequences.
The Ultimate Benchmark: A Believable Clean
There was one thing that, no matter how hard I tried, none of the available models at the time could get right: a believable clip of a woman brushing her teeth.

It sounds simple. Yet, in my early short films, I had to actively hide or cut around this action to hide the digital hallucinations. Eventually, I realised this mundane human activity made for a perfect benchmark of progress; so I resolved that with the release of every new video model, I’d try the test again.
And with every subsequent model release, the test failed again. Sometimes in weird ways, sometimes in hilarious ways, sometimes in disturbing ways. But it always failed.
Photorealism, Speech, and the Breakout (2025)
Throughout 2025 the next generation of video models, such as Hailuo 2 and Wan, brought higher detail, improved motion, precise camera control, and believable near-photorealism. None of them could pass my test.

Google's Veo 3.0 was the first video model to natively generate synchronized sound and speech. This was a true breakout moment for AI video, notably manifested in the highly talked-about Kalshi advertisement played on US TV during the NBA finals.
But even on Veo 3.0, my toothbrushing test failed. We could do so much, but some things stayed out of reach. This was hard to explain to clients who’d only ever seen the finished results, not the work in progress.
At the end of 2025, we started getting even more powerful and capable models. With the right prompting, you could get them to look incredibly, believably "real." But they still couldn't depict a believable tooth brushing. Each new model got closer and closer, but they just couldn't get there.
The Benchmark is Passed (2026)
Come early 2026, a new wave of video models arrived, starting with Kling 3, and then the big one: the current best-in-class, Seedance 2.
And what had seemed impossible, happened. My toothbrushing test was passed. Not flawlessly, not perfectly consistently, but enough to be more than a fluke.

It's incredible to look back at how far we've come in just three years; from the flickering, low-motion and unstable four-second clips of Runway Gen-2 to the jaw-dropping realism and long multi-scene sequences of Seedance 2.
The latest models can generate longer, highly coherent clips at up to 4K quality, complete with native sound, flawless lip-syncing, and precise editing capabilities. New tools make the act of creation more controllable and straightforward.
That incredibly complex scene we worked so hard to execute in our first film, finding faith we can now make in a fraction of the time, using just a single reference image and Seedance 2. What once took weeks, now takes days; what took days, now takes hours. This is going to be made even more efficient by adopting increasingly capable agentic workflows.
The Elevation of Craft
This rapid evolution has democratised production. AI video now enables smaller brands to punch far above their weight, allowing them to produce ads that would once have been beyond the reach of their production budgets.
But as the barrier to creation is lowered, the importance of craft is elevated.
Some of the more obvious technical "tells" of AI video—the flickering, the morphing, the floating objects—have largely disappeared. But this puts more scrutiny on the actual artistic performance.
Perhaps I'm speaking with the curse of knowledge, but when looking at recent high-profile campaigns I can see that Secret Escapes, for example, looks clearly AI-generated — the smoothness and symmetry are giveaways — while with this OMHU ad it’s much harder (if not impossible) to say the same.
The difference isn't the software; it's craft and discernment. With an abundance of visual options now available, the creative question of why you picked a specific artistic solution becomes more important than ever.
The Next Frontier: World Models
No matter how good these generative models get, there are always more edge cases waiting to be found. My toothbrushing test may have been passed, but these are a few new examples I’ve come across:
- A man counts from one to five on his fingers.
- One person kicks another, causing them to fly backwards and hit a wall.
- A falling wine glass spills red wine onto a rug, then smashes.
These scenarios are easy enough for even a child to comprehend, but current video models struggle immensely with them. This is because today's models are fundamentally visual mimics. They don't actually understand cause and effect; they don't understand complex, chronological sequences of events, and they don't understand the physical rules of gravity, momentum, and impact.
Enter the future of video: World Models, like DeepMind's Genie 3.
These models are being trained on vast, multimodal datasets so they can begin to understand the physical rules of our world—its geometry, physical relationships, and complex interactions.
We're still in the absolute infancy of this transition. And when they finally mature, I'm sure we'll find a whole new set of limitations to work around. But as creators, we should welcome that. After all, it's limitations that provoke true innovation.