Hey folks — PhD student at the University of Maryland here, studying how developers debug and iterate on multi-agent systems. We’re in the last stretch of recruitment, with sessions running now through next week. The question we’re testing: when you tweak a prompt in an agent workflow, you usually judge it by eyeballing a run or two. Our research tool shows the distribution of outputs each node produces across runs — does that actually beat clicking through traces one at a time, or is it just one more dashboard? “It doesn’t help” is a publishable answer. Participating: a 75-min Zoom session on structured debugging tasks (recorded, think-aloud), about a week using the tool in your own workflow, and a 30-min follow-up interview. $150 gift card on completing the full study. If you’ve built with LangGraph/LangChain (or agent workflows generally), the screener takes ~2 min: https://forms.gle/Zwqvgd1h8DUnFRfC8 IRB-approved academic research, not a product pitch. Questions welcome — or zxu169@umd.edu . submitted by /u/LeoXzz
Originally posted by u/LeoXzz on r/ArtificialInteligence
