Original Reddit post

What it is. video-lens is a free, MIT-licensed Claude Code skill for macOS. Claude can’t watch video, so video skills hand it frames. video-lens measures every frame first (UI motion timing and easing as CSS, scene cuts, on-screen text, speech with timestamps) and gives Claude numbers and text, with images only where a look is needed. Nothing is uploaded: it runs on ffmpeg, OpenCV, macOS Vision and on-device speech recognition. How Claude helped. I built it with Claude Code on Opus 5.5: the spec, the Python analysis code and Swift helpers, 60 self-tests on synthetic clips with known answers, the benchmark harness and the scorer. My part was deciding what to measure and checking the results against answer keys. Benchmark. 9 tasks scored against answer keys (CSS animations rendered in headless Chrome, Korean lectures, a real screen recording), 27 runs per condition, Claude Opus 5.5, one run at a time: What I learned

  • On a strong model the skill barely moves accuracy (0.949 to 0.956). What it changes is cost, time and how bad the worst run gets. On a weaker model the gap is bigger: Grok 4.7 went from 0.819 to 0.942. - With /watch and video-use, Opus usually also wrote its own ffmpeg or OpenCV analysis on top of the skill (23 and 27 of 27 runs, against 2 of 27 with video-lens). That is where most of their extra cost and time went. - I had Opus rebuild two viral launch-week showreels as HTML from the video alone. Without the skill the rebuilds looked slightly closer to the originals (SSIM 0.82 and 0.84 against 0.79 and 0.81); with it they cost 24% less. The side-by-side videos are on the site. - macOS only, Python 3.10 or newer. Try it (free) /plugin marketplace add Junhan2/video-lens /plugin install video-lens@video-lens Repo: https://github.com/Junhan2/video-lens Results, raw data and the scripts to rerun them: https://junhan2.github.io/video-lens/ Feedback on where it breaks is very welcome. submitted by /u/superpuppycrawl

Originally posted by u/superpuppycrawl on r/ClaudeCode