01In plain English
A benchmark is a fixed set of tasks, such as maths problems, coding challenges or exam questions, used to score AI models against each other. Model makers publish benchmark results when they launch.
02Why it matters when you are choosing
Benchmarks show roughly how capable a model is, but a high score on a test does not guarantee good results on your work. Some models are tuned to do well on the popular tests.
03What to check
Ask the vendor
Treat benchmark claims as a starting point and try the tool on your own tasks before you commit.