GPQA.

A graduate-level, 'Google-proof' multiple-choice benchmark of hard science questions designed to resist easy lookup and test deep reasoning.

GPQA (Graduate-Level Google-Proof Q&A) is a challenging benchmark consisting of multiple-choice questions in biology, chemistry, and physics at the graduate level. Unlike many AI benchmarks that can be gamed through memorization or simple web searches, GPQA questions are specifically designed to require deep scientific reasoning and understanding. The benchmark serves as a critical test for evaluating whether AI models truly comprehend complex scientific concepts rather than simply retrieving information.

The 'Google-proof' aspect means these questions cannot be easily answered by searching online, as they often require synthesizing multiple concepts, applying theoretical knowledge to novel scenarios, or reasoning through multi-step scientific problems. Questions are crafted by domain experts and validated to ensure they test genuine understanding rather than factual recall. The benchmark typically includes detailed explanations for correct answers, allowing for nuanced evaluation of model reasoning processes beyond simple accuracy metrics.

GPQA has become particularly important for evaluating reasoning models and advanced AI systems, as it reveals gaps between surface-level knowledge and true scientific understanding. While top AI models have made significant progress on GPQA, perfect performance remains elusive, highlighting the continued challenge of achieving human-level scientific reasoning. The benchmark's resistance to gaming makes it a reliable indicator of genuine AI capabilities, though critics note that graduate-level questions may not fully capture the breadth of scientific reasoning required in real-world applications.