Original Reddit post

Not quite for this channel, but relevant to choosing new-age apps: users don’t pick the best feature; they pick what they know. I built a simple bandit model to test whether usage converges on quality over time. It doesn’t estimate which option is better. It simply reinforces whatever gets used more. Three update rules produce very different outcomes: one locks onto early winners regardless of quality; one mostly self-corrects but can still get stuck under strong reinforcement; and one control always finds the true best option. Same mechanism, wildly different UX. Basically, habit formation in miniature. Feedback welcome :) submitted by /u/svk_roy

Originally posted by u/svk_roy on r/ArtificialInteligence