Original Reddit post

i have been thinking about what happens when we spend more compute checking a model’s output. it can give us finer scores, more consistent judgment, and clearer reasoning but it cannot give the verifier evidence it never had. for example one verifier might call both 20 and 24 wrong when the answer is 25. another can tell whether 24 is closer and useful but it still does not tell us whether the verifier is checking the right thing. i thought about writing a short note about it here. curious what others think? https://www.mindmodelmachines.com/notes/what-more-verifier-compute-actually-buys submitted by /u/svk_roy

Originally posted by u/svk_roy on r/ArtificialInteligence