Public Benchmarks Unreliable Due to Data Contamination
Publicly published AI benchmarks are increasingly unreliable because many are now included in the training datasets for new models, according to Nathaniel. This integration means models are trained on the very metrics intended to objectively evaluate them, diminishing their value.
Frontier AI models are being released at an average rate of one every 11 days, making it difficult for benchmarks to remain current and unbiased.
While companies are not intentionally misrepresenting results, the continuous inclusion of benchmarks in training data compromises their ability to serve as objective measures of performance over time.


