Original Reddit post

There’s a benchmark called MedXpertQA that has difficult board level medical questions and has the questions organized by organ system. I seem to be only one to have run this benchmark against the latest models including from OpenAI/Google and published the results online. 2 years ago GPT-4o scored 42.8% and now Astra is approaching saturation at 87.4%. You can see the results and filter by organ system here . Let me know if you have any questions or want me to run against any other models. submitted by /u/toshv

Originally posted by u/toshv on r/ArtificialInteligence