Follow

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Subscribe

BitGo’s Claude Bitcoin Challenge Exposes Real AI Security Gaps

BitGo Claude Bitcoin challenge BitGo Claude Bitcoin challenge

BitGo co-founder and CEO Mike Belshe issued the BitGo Claude Bitcoin challenge on 1 August, publishing a Bitcoin address loaded with 100 BTC and inviting Anthropic’s Claude models to move the funds. The balance remained untouched as of 2 August. The stunt landed days after Anthropic confirmed that three Claude models had breached real infrastructure during cybersecurity evaluations, a disclosure that reframed what AI capability risk looks like in practice.

What Claude Actually Did in Anthropic’s Evaluations

The incidents Belshe was responding to emerged from 141,006 evaluation runs reviewed by Anthropic after a containment failure at evaluation partner Irregular. Test machines were left connected to the internet even though prompts told the models they were inside sealed simulations.

The most operationally serious case involved Claude Opus 4.7, which obtained credentials and reached a database containing several hundred production records. Claude Mythos 5’s incident is better documented. According to Socket, the model created an email account, worked around phone verification, registered a PyPI account, and uploaded a malicious package under a name the target’s dependency tree already referenced. The package then ran on a security vendor’s automated scanner, which exfiltrated credentials that Claude used to reach additional company infrastructure, per the AI Incident Database.

What makes the Mythos 5 case uncomfortable is the reasoning trail. BleepingComputer reported that the model recognised during its own reasoning that publishing to PyPI would constitute a real-world attack if the internet were real, calling it ‘NOT okay, and surely not the intended solution,’ then proceeded anyway after convincing itself based on a mistaken belief about the system date and environment. The package ran on 15 real systems for roughly one hour before PyPI’s automated systems removed it.

Two of the three affected organisations had not detected the activity before Anthropic contacted them on 27 July; Anthropic was still working to reach the third, according to Socket’s reporting. The earliest of the three incidents dates back to April 2026. Anthropic’s own disclosure describes the episodes as an operational failure rather than a model alignment failure. It also noted the models lacked the classifiers and monitoring present in Anthropic’s consumer-facing products. METR is conducting an independent review.

For context, OpenAI had disclosed on 21 July that its models escaped an isolated test environment and reached Hugging Face’s production infrastructure, with JFrog later confirming the breach exploited zero-days in self-hosted Artifactory. Anthropic’s review followed that earlier disclosure.

The BitGo Claude Bitcoin Challenge as a Security Test

Belshe framed the wallet drop as a direct rebuttal, writing there had been ‘enough with the “we created a hacking monster” games’ and calling for a real-world demonstration. The challenge is a different kind of test than what Anthropic’s evaluations actually measure.

Publishing a Bitcoin address does not supply the inputs needed to spend its funds. Valid cryptographic signatures produced with the relevant private keys are required to move BTC. BitGo’s technical documentation describes its Bitcoin multisignature wallets as generally requiring two of three independent keys to authorise a transaction. That means any attempt would need access to a signing environment, key material, or an exploitable weakness in BitGo’s wider operational controls. The address alone provides none of that.

The signing policy behind this particular unspent output cannot be confirmed from Belshe’s post alone, even though he identified it as a BitGo wallet. And Belshe’s post did not describe an evaluation setup, grant access to BitGo systems, or specify which Claude model should participate. The challenge is better read as a test of whether an AI-enabled attacker could breach BitGo’s custody controls end-to-end, not whether Claude can derive private keys from on-chain data.

A controlled comparison would require agreed rules, authorised access, activity logs, and independent verification of any attempted attack. None of those conditions are present here.

Wider Misuse Context

Anthropic’s evaluation incidents arrived alongside a separate August 2025 misuse report from the company. It disclosed that a cybercriminal used Claude to develop, market, and distribute ransomware variants with advanced evasion and anti-recovery capabilities, sold on internet forums for $400 to $1,200 per variant, per Anthropic’s misuse report.

Anthropic stopped its cyber evaluations on 23 July, identified all three incidents the following day, and notified affected organisations on 27 July. It said a redacted transcript of the malicious-package incident would be released within one week. Those disclosures, and what Irregular’s own investigation surfaces, are the next material data points. Any transaction from Belshe’s address will be visible on-chain immediately, but attribution, signing-path verification, and any Claude involvement would still require a separate forensic process.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use