Original Reddit post

I always thought this could be a good proxy on model’s intelligence, so decided to try a minimal test: models write a single prompt template for gpt-2, and gpt-2+the template is scored on 395 examples of a basic task (deciding the correct action for a farm give a short report). Admittedly limited usefulness as a traditional benchmark, but I wrote a short write up that I think reveals some interesting points about model capabilities. submitted by /u/pokeuser61

Originally posted by u/pokeuser61 on r/ArtificialInteligence