About a year ago I started experimenting with what I thought was a fairly simple question: can current generative video actually sustain a long-form narrative rather than a trailer or a collection of impressive shots? I expected to make one short film. I ended up making seven films of roughly an hour each. Here you have the movies: https://youtube.com/@jordibenach The project became an accidental longitudinal experiment in AI filmmaking because the technology was changing while I was making it. I started with Sora 2 when it had only been available for about a month. C1-C3 were assembled in iMovie on my iPhone, sometimes during lunch or while watching TV. By C4 I had moved to Final Cut Pro, and later films use a much more elaborate workflow involving different AI systems. What surprised me most is that generating beautiful individual clips isn’t really the hard problem. The hard problem is production memory. After hundreds or thousands of generations, somebody has to remember what every character looks like, what they’re wearing, where they are, what happened previously, how old they should be, and what the emotional history between them is. Eventually I stopped expecting the model to remember the production. I became the production memory and gave the model much smaller, controlled problems to solve. Another thing I learned is that editing creates complexity. The individual generation doesn’t have to. The early films are rough, and I’ve deliberately left them that way. C1-C7 now document two things running in parallel: the story in the films, and me learning how to make them while the technology itself was evolving. I’m curious whether other people experimenting with long-form generative video have reached the same conclusion: are we moving from a prompting problem to a production problem? Disclosure: this is my own non-commercial film project. submitted by /u/bdylsing
Originally posted by u/bdylsing on r/ArtificialInteligence
