โ† Working Theses

A Modern Odyssey ๐Ÿ›๏ธ

A modern retelling of The Odyssey, Poseidon becomes an AI system, the Cyclops becomes a security sensor, the Sirens become AI-reconstructed memories. Every frame generated with AI: Google Flow, Kling 3.0, Hailuo, Vidu, Seedance 2.0, and ComfyUI.

It started in a car. My son and I had just watched The Odyssey in the theater, and on the drive back we got talking about what it would actually mean to lose all communication with home and not know your way back. No coordinates, no signal, no way to tell anyone you're alive. That conversation stuck with me longer than the movie did.

Around the same time, I'd been curious about Google Flow since it was announced at Google I/O, the kind of curiosity that sits in the back of your head until you finally clear a weekend (which ended up being multiple weekends) to actually try the thing. Those two threads collided into one idea: build a modern Odyssey, and build it entirely with AI.

The result is a four-minute sci-fi short called The Long Way Home, swapping Poseidon for an autonomous AI system called ATHENA, the Cyclops for a single-lens security sensor, and the Sirens for AI-reconstructed memories of the people you love. 40+ generated clips, six different video models, and one very persistent content filter later, here's what I learned.

๐Ÿ›๏ธ Why the Odyssey

I didn't want to write a sci-fi story and bolt Greek mythology onto it for flavor. I wanted the structure to carry the meaning, the Trojan Horse as a maintenance drone that smuggles you past a security perimeter, the temptation of Calypso's island as a community that's simply stopped trying to leave, the Sirens as an AI that can almost, but not quite, reconstruct the people you're trying to get back to.

The myth isn't decoration here. It's the reason the story works at all, every obstacle Matt faces isn't really trying to kill him. It's offering him a reason to stop.

๐ŸŽฌ The tool stack, and why it kept changing

Google Flow was where I started, and where I spent the most time fighting the gap between what I expected and how it actually works. If you've spent any time with an LLM, you expect context to carry forward, you correct one thing, and the next response remembers it. Flow doesn't work that way. Every generation is closer to a blank slate, which means building a scene frame by frame takes real, deliberate effort. Putting together even a rough storyboard eats time you don't expect to spend, and accuracy isn't reliably there even after several attempts at the same prompt.

The pricing is its own puzzle. Google's free tier genuinely goes further than most of the other tools I tried, but "further than the others" still isn't far enough to finish a project on. And there's no way to simply buy more credits at the lower tiers when you run out; your only option is to move up a full subscription tier, which feels like a gap in an otherwise thoughtful product. Flow does have an extend-scene feature and a way to copy a prior prompt forward, which sound like exactly what you'd want for continuity, but neither works especially well in practice, because the context behind a shot still isn't something the tool actually holds onto.

On top of that, Flow's "prominent people" filter flagged a completely fictional AI-generated face, twice, for reasons Google never fully explained. I lost real production time to that before realizing the fix was regenerating the reference image with more distinct features, not rewriting the prompt.

Kling 3.0 became my go-to for short scenes. The detail in individual shots is genuinely better than Flow's, sharper textures, more convincing motion on things like the Cyclops's mechanical arm. But it doesn't flow seamlessly with the rest of the stack. Moving a shot from Kling into a sequence built around Flow's look and pacing takes real massaging; the two tools just don't think about a scene the same way.

Hailuo earned a place in the rotation for anything needing smooth, simple camera motion and zero account friction. Good for connective tissue shots I didn't want to burn a "real" credit on.

Vidu was my answer to the single hardest problem in this entire project: keeping Matt's face consistent across forty separate generations. It didn't solve it completely, but it solved it better than anything else I tried.

Seedance 2.0 is good for standalone scenes, it creates strong continuations from a single starting frame, and the individual clips hold together well on their own. But for longer runs, where I needed a sequence of shots to feel like one continuous story rather than a series of impressive individual moments, I kept switching back to Google Flow. Seedance wins the single shot; Flow still wins the sequence.

ComfyUI is worth it if you're expecting a genuinely large amount of work, a full pipeline you'll reuse across many shots. But it demands a lot more from your system than any of the hosted tools, and unless you're doing this professionally, it's overkill for a project this size. Full control over the generation pipeline, node by node, at the cost of a steep learning curve that ate an entire evening before I produced anything usable.

ElevenLabs handled every voice in the film, including ATHENA, whose "not quite human" quality turned out to be a post-processing decision, not a generation one. Pitch down two semitones, add a faint delay on sentence-final words, cut the frequency range like a bad phone line. The AI never had to sound cold. I just had to make it sound cold afterward.

Clipchamp did the actual assembly. Free, browser-based, no watermark, and genuinely surprising, good enough auto-captions that I only had to hand-correct the distorted ATHENA lines.

๐Ÿšง What actually broke

Character consistency is the real cost of "free." Every model treats each generation as stateless. Matt's face, jacket, and even the facility's geometry drifted shot to shot unless I anchored every single generation to a locked reference image. This is the unglamorous, unskippable work of AI filmmaking, building a character bible before you touch a single video prompt.

Content filters don't explain themselves. I hit two different flags on the same platform for two different reasons, one for a prompt that implied a character's death too literally ("fire," "falls to the ground," "fades into darkness" reads as a death scene even when you mean something more ambiguous), and one for a reference image the model's likeness-matching decided looked too close to a real person. Neither came with a way to actually understand why. The fix, both times, was the same: change what you're asking for, not how you're asking for it.

Time and aging are harder than action. I wanted one continuous camera orbit around Matt, starting young, ending ten years older, no cuts, using the camera's own rotation to hide the transformation. It's a real cinematography technique. It's also right at the edge of what any current model can hold in a single generation. I ended up planning both the single-take version and a two-clip fallback with a whip-pan disguising the cut, because betting the whole shot on one model succeeding wasn't worth the risk.

Nobody has built the "understand my scenes" tool yet. I have over forty ten-second clips that don't line up cleanly, and every AI editing tool I found is good at technical continuity, matching lighting, extending a shot by two seconds, fixing a flicker and completely uninterested in narrative continuity. Whether Matt's jacket matches. Whether the emotional beat lands. That part is still entirely on me, and I don't think that's changing anytime soon.

๐Ÿงญ The actual thesis

No single tool gets you home. Flow for the sequences that need to feel continuous, even with its friction. Kling when a single shot needs more detail than Flow can give it. Hailuo for the connective shots that don't need to be perfect. Vidu when a face has to survive forty generations. Seedance when a scene stands entirely on its own. ComfyUI when the scope of the work justifies the system cost and the learning curve.

That's not a compromise. That's the actual skill. Knowing which tool gets you through which chapter and being willing to switch mid-sentence when one of them says no.

Ten years, one raft, and a smartwatch got Matt home. It took me a stack of six AI models and a lot of regenerated reference images to get the story told. Odysseus didn't have it any easier