A volunteer group called the Bitcoin Red Team just ran one of the most ambitious automated security audits the crypto ecosystem has ever seen. Their weapon of choice: Kimi K3, an open-weight AI model built by China’s Moonshot AI. Over roughly 108 hours, the model catalogued 7,958 potential security findings across 501 Bitcoin-related open-source projects, with 1,280 of those rated high or critical severity.
Kimi K3 outperformed every other open-weight model tested, including Zhipu’s GLM-5.2, in standardized vulnerability detection benchmarks.
What the audit actually found
Of the 7,958 potential issues flagged by Kimi K3, only 24.7% could be dynamically reproduced. At the time of reporting, 29.4% of the findings had been communicated upstream to the affected projects.
The most consequential discovery was a critical two-factor authentication bypass in BTCPay Server version 2.4.2. The vulnerability had already been exploited to extract Lightning wallet credentials before it was patched, making it a live, in-the-wild security incident rather than a theoretical concern.
How Kimi K3 stacks up
The UK’s AI Safety Institute and its counterpart CAISI ran a preliminary assessment of Kimi K3 in July 2026, scoring it at 32% on ExploitBench. That’s a benchmark designed to measure an AI model’s ability to identify and reason about exploitable software vulnerabilities. GLM-5.2 scored 24% on the same test.
Among open-weight models, those whose weights are publicly available for anyone to download and run, Kimi K3 sits at the top. The model was released around July 16–27, 2026, and the Red Team intensified its auditing effort in the weeks that followed.
The gap between open-weight and closed-source models remains significant. Leading US models from OpenAI and Anthropic averaged around 76% on ExploitBench. That’s more than double Kimi K3’s score.
The Coldcard incident that started it all
The Red Team’s effort was catalyzed by a security incident in July 2026 involving the Coldcard Mk3. A flaw in the Mk3 firmware led to the theft of approximately 594 BTC, estimated at $38 million at the time. The incident sparked widespread speculation that the attackers had used AI to identify the firmware vulnerability, though that claim hasn’t been definitively proven.
What this means for Bitcoin security
The economics of code auditing are about to shift. A professional security audit of a single Bitcoin project can cost tens of thousands of dollars and take weeks. Kimi K3 scanned 501 projects in 108 hours.
US AI companies, which build the most capable models, have generally restricted their tools from being used for vulnerability research, citing safety concerns. Meanwhile, an open-weight Chinese model is being freely deployed to find and report bugs in critical financial infrastructure. The gap between open-weight and closed-source model performance on ExploitBench—32% versus 76%—suggests the most capable vulnerability detection still lives behind API paywalls.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
22









English (US) ·