Pinned post

A look at the more challenging AI evaluations emerging in response to the rapid progress of models, including FrontierMath, Humanity's Last Exam, and RE-Bench (Tharin Pillay/Time)

Tharin Pillay / Time : A look at the more challenging AI evaluations emerging in response to the rapid progress of models, including Fron...