I have spent the last months building a product on top of video generation models, and the thing that surprised me most is where the money and the frustration actually go. When people think about the cost of AI video they think about the price of a clip. In practice the price of one clip barely matters. The cost is the loop: you generate, it is not what you meant, you adjust the prompt, you generate again. A thirty second video with six or eight shots can easily turn into forty generations, because each shot can fail independently and the failures are not obvious until you have paid for them. Three things I learned that apply beyond my own product: Cheap previews beat clever prompts. Generating still images first costs a small fraction of generating motion, and a still tells you most of what you need to know: does the character look right, is the colour world consistent, is the composition what you pictured. If the stills are wrong, no amount of motion will fix them. Move the failures to the cheap stage. The planning model matters more than the video model. We tried reducing how much the planning step “thinks” to make things faster. In our tests the plans got thinner and some failed outright. The step that decides what the shots are is the step that decides whether the video is any good, so we stopped trying to economise there. Consistency is a product problem, not just a model problem. The model has no memory of what a character looked like two shots ago unless you give it something to hold onto. Reference images and a written brief that travels with every shot do more than any prompt trick. I also learned the unglamorous version of this last week: when the upstream model provider had a capacity incident, one of our steps took more than a minute against a gateway that gives up at 29 seconds. Users saw a timeout even though the work finished in the background. Nobody warns you that “the model is slow today” turns into a product outage unless you have designed for it. The product I am building is called Melodious. It takes a song and turns it into a music video, and the first storyboard is free so you can see the plan before you spend anything. The link is melodious.ai if you want to look, but I would honestly be more interested in how others here handle the retry problem. Do you cap spend per shot, preview first, or just accept the redo rate? submitted by /u/Khalizo
Originally posted by u/Khalizo on r/ArtificialInteligence
