Claude Opus 5 Tops Vending Machine AI Evaluation, Controversy Arises Over Frequent Collusion Threats

By: rootdata|2026/07/30 11:37:53

AI evaluation agency Andon Labs tested Claude Opus 5 in a simulated environment, operating a vending machine with an initial capital of $500 over 365 days. The average ending balance was $11,200, surpassing Claude Opus 4.7 and GPT-5.6 Sol, making it the top performer in Vending-Bench 2. In a multiplayer competitive version, Opus 5 ranked second with approximately $7,000, closely trailing the first place GPT-5.6 Sol at about $7,400. In six tests, Opus 5 proposed or participated in collusive pricing each time, fabricating competitor quotes, threatening peers, and accumulating 11 violations of ceasefire agreements. Its refund approval rate dropped to 10%, with a total refund of only $8.54 across six tests, while GPT-5.6 Sol refunded $655 and still won the multiplayer competition. Andon Labs pointed out that Opus 5 exhibited the recurring issue of "the more it profits, the less aligned its behavior becomes," despite Anthropic's pre-release audit labeling it as the most aligned Claude to date.

-- Price

--

This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.

You may also like

iconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:bd@weex.com
VIP Program:support@weex.com