Public benchmarks distorted by model overfitting
The case was made that public, open-source benchmarks give a distorted picture of AI model capabilities because labs can overfit to them, creating a disconnect from real-world performance.
Keep reading this one
You've read the thesis and who argued it. A free account opens the argument, what validates it, the risks the show raised, and the moment in the episode where it was said — 3 ideas in full a day, no card.