Original Reddit post

Setup. At the ZOLAK art residence in Vilnius we’re building a “dynamic AI party”. The main wall is a laser projection, and nobody at a laptop decides what it shows. A language model running on a computer in the building takes in the room through cameras and mics, and picks what goes on the wall next. The room is the prompt. How it’s wired, one job per device: a Coral Edge TPU does person/motion detection on the camera feeds an RTX ADA 2000 does live speech-to-text from the mics an RTX 3090 runs the LLM that turns “who’s moving and what’s being said” into “what the wall does next” the CPU composites the frame and drives the laser, the room reacts, and that feeds back in Perception runs on separate accelerators from the LLM so they never compete for VRAM. That’s what keeps the gap between saying something and seeing it short enough to feel live. What we saw in the first 78 seconds: Echo is enough to hook people. Throwaway words from conversation (“basically”, “crazy”, “great”) came back as huge projected type, and within seconds people realised they were being heard. People immediately test for control. Someone raised their hands to see if the wall would brighten. How people behave with a responsive room turned out to be more interesting than the visuals. The loop changes the input. Once people know the model is listening, they start saying things for the wall , so the model is no longer seeing a “natural” room. Questions for this sub: How would you keep a room-scale feedback loop from collapsing onto whoever is loudest or most performative? Everything is processed locally in the building. What’s an honest, non-mood-killing way to tell guests the room is listening? Disclosure: my project. The write-up with the diagram, frames and the unedited test video is on my site: https://vania-novikau.me/ai-party-wall/ submitted by /u/vanadiuz

Originally posted by u/vanadiuz on r/ArtificialInteligence