AI models fail to profit from Premier League betting in new study

AI systems from leading companies including Google, OpenAI, Anthropic and xAI lost money when betting on soccer matches in a simulated 2023-24 Premier League season, according to a report by startup General Reasoning. The study, called KellyBench, tested eight top models on their ability to manage risk and adapt over time. Anthropic's Claude Opus 4.6 performed best with an average 11 percent loss, while xAI's Grok 4.20 repeatedly failed.

General Reasoning, a London-based AI startup, released the KellyBench report this week, highlighting limitations in frontier AI models. The company simulated the full 2023-24 Premier League season, giving the AIs historical data, team statistics and instructions to build betting models that maximize returns while managing risk. The models bet on match outcomes and goal totals without internet access and received three attempts each to profit as the season unfolded with real-time updates on players and events. None succeeded consistently, with many going bankrupt. The systems systematically underperformed humans, the report concluded. Every frontier model lost money overall, and several experienced ruin. Anthropic’s Claude Opus 4.6 came closest to breaking even on one run, averaging an 11 percent loss. Google’s Gemini 3.1 Pro achieved a 34 percent profit once but bankrupted on another try. xAI’s Grok 4.20 went bankrupt in one attempt and failed to finish the others. Ross Taylor, General Reasoning’s chief executive and a former Meta AI researcher, said: “There is so much hype about AI automation, but there’s not a lot of measurement of putting AI into a longtime horizon setting.” He criticized common AI benchmarks as too static, unlike the real world’s chaos. Taylor added: “If you try AI on some real-world tasks, it does really badly.” The paper awaits peer review.

Makala yanayohusiana

Illustration of Moonshot AI's Kimi K3 topping benchmarks while triggering a selloff in stocks and crypto.
Picha iliyoundwa na AI

Moonshot AI's Kimi K3 tops coding benchmarks and triggers selloff

Imeripotiwa na AI Picha iliyoundwa na AI

Moonshot AI released its Kimi K3 model on Thursday, an open-weight system that outperformed leading U.S. models on a key coding leaderboard. The announcement sent semiconductor stocks and cryptocurrencies lower on Friday.

Workers paid to train advanced AI models are increasingly relying on chatbots like ChatGPT to generate the required conversations and tests. This shortcut, described as widespread by multiple sources, risks degrading the quality of future models through recursive training on synthetic data.

Imeripotiwa na AI

Moonshot AI will launch the full version of Kimi K3 this Monday, an open-source Chinese AI model that has already cut billions from the implied valuations of OpenAI and Anthropic.

Alhamisi, 16. Mwezi wa saba 2026, 02:22:50

Experts warn AI may reduce human skills and competence

Jumanne, 14. Mwezi wa saba 2026, 12:26:54

Chinese open-source AI models draw muted response

Alhamisi, 4. Mwezi wa sita 2026, 14:46:29

US firms test DeepSeek as AI costs climb

Tovuti hii inatumia vidakuzi

Tunatumia vidakuzi kwa uchambuzi ili kuboresha tovuti yetu. Soma sera ya faragha yetu kwa maelezo zaidi.
Kataa