What is an AI benchmark?
News stories often say a new AI model "sets record scores" on benchmarks. What does that actually mean?
A benchmark is, put simply, an exam for an AI model. It's a fixed set of questions or tasks — for example, math problems, coding exercises, or reading comprehension questions — used by researchers to measure how well a model performs. Because every model gets the same questions, the scores can be used to fairly compare different models against each other.
So when you read in the news that a new model "sets a record score on a benchmark," it means it scored better on that specific exam than previous models. That's useful information, but with a catch: a high score on one benchmark doesn't automatically mean a model feels better or is more useful in practice. Models can be specifically trained to do well on exactly that kind of exam question, without that always translating into better overall performance.
Compare it to a student who becomes very good at doing practice tests from one particular method, without that guaranteeing they truly master the material broadly. For that reason, serious AI researchers always look at multiple benchmarks at once, and increasingly also at how a model performs in real-world use, not just on paper.
For you as a reader: take a single "record score" headline with a grain of salt. It tells you something, but not the whole story.