Frontier models ship monthly; team-context performance shifts with them. This living leaderboard scores each model on grounded tasks — retrieval-heavy Q&A, agent tool-use, long-context synthesis — and is updated as new models land.