Original Reddit post

Official model page: https://seed.bytedance.com/en/seedance2_5 LingBot-Video paper: https://arxiv.org/abs/2607.07675 The 30-second limit on ByteDance’s Seedance 2.5 page sent my mind to a kitchen timer, not a cinematic prompt. At that length, a model has time to lose track of something small it showed near the start. The page also lists reference control and editing. The first test I want to run is almost embarrassingly plain. Put a digital kitchen timer at 20 seconds and lock the camera. Someone folds a tea towel once, sets it beside the timer, and waits. At zero, the timer beeps and the person presses stop. The towel gives the model a second state to hold onto while the display keeps changing. I have not run this yet. I would count a pass only if the numbers move down in order, the beep lands at zero, the hand stops the timer afterward, and the folded towel does not reset halfway through. The digits create a problem of their own. A badly drawn 8 is not the same failure as a countdown that stalls or runs backward, so I would score those separately. Reference control makes the result harder to interpret. I would use the same prompt and whatever other settings the interface exposes, first with minimal reference material and then with much more guidance. I would repeat both conditions a few times, since one clean generation could just be luck. If the guided clips fail less often across those repeats, that would be evidence that the references helped. I still could not say the model handled the full sequence without that support. I would put LingBot-Video through the same scene. Its paper describes rewards for physical rationality and task completion, so the timer, the handoff and the towel state fit the behavior I want to measure. I would use the same score sheet for both models and skip a general quality ranking. Both models may handle this without trouble. If they do, I would move the timer partly out of frame for a few seconds and check whether it returns at the right point in the countdown. If a clip breaks, the separate scores should show whether the problem came from the digits, elapsed time, sound or object state. submitted by /u/Brave_Pressure_9886

Originally posted by u/Brave_Pressure_9886 on r/ArtificialInteligence