Every score is worked out from the vendor's own pages · How we score →Disclosure
SAASINSPECTOR
AI concepts

Benchmark

A standard test used to compare AI models on the same set of tasks.

01In plain English

A benchmark is a fixed set of tasks, such as maths problems, coding challenges or exam questions, used to score AI models against each other. Model makers publish benchmark results when they launch.

02Why it matters when you are choosing

Benchmarks show roughly how capable a model is, but a high score on a test does not guarantee good results on your work. Some models are tuned to do well on the popular tests.

03What to check

Ask the vendor

Treat benchmark claims as a starting point and try the tool on your own tasks before you commit.

← Back to the glossary