Bitcoin Red Team, an AI-assisted security campaign targeting Bitcoin projects, reports 6,700 findings across 425 projects during its first 55 hours. Of those, 1,029 carried a high or critical label.
The number it hasn’t given out is how many held up.
That omission is the story. What the Aug. 6 update actually measures is the volume of material fed into a triage pipeline — not the amount of software that ended up safer. The campaign’s security impact goes unreported.
The missing numbers are the ones that matter
Nowhere in the thread are there audit-ready definitions for the severity labels, or the denominators sitting behind them. No case-level outcomes. No aggregate false-positive rate. No fix rate.
Strip those fields out and there’s no way to work out how many alerts hardened into confirmed vulnerabilities, how many maintainers threw out or downgraded, or how many ended in a patch. All that’s left is a headline count and a severity split — and both are self-assessed.
None of which renders the 55 hours pointless. It demonstrates something worth taking seriously: an AI system can flood a review pipeline at the scale of an entire project set, and do it fast. Everything after that stage — expert prompting, reproduction, disclosure, maintainer response — still had to happen by hand.
Two snapshots, and what changed between them
Progress was published twice. The 27.5-hour mark showed 390 projects and 4,962 findings. By the 55-hour mark, projects were up by 35 and findings by 1,738.
High-or-critical findings came to 15.4% of the total in the later thread, which also clarified that three of the 24 reported participants were bots.

The way the accounting shifted is worth a second look: critical and high were reported separately in the earlier post and merged in the later one. Campaign assessments underpin both sets of figures. Establishing maintainer-confirmed exploitability and remediation outcomes takes separate evidence, and neither post contains it.
The models searched. People decided.
Per Rob Hamilton, Kimi K3 carried the heavy analysis, while GPT Sol, Fable/Opus and GLM 5.2 supported the documentation side. Selected components he considered load-bearing were covered by OpenAI’s Cyber Harness, he said.
Hamilton wrote a day later that subject-matter experts could flip an assessment with one or two sentences of context or a small block of code. In the cases he described, that input lifted middling concerns into high or critical territory.
He was specific about the bottlenecks too: operations, disclosure handoff and triage. Model capacity wasn’t on the list.
By his account, the models did broad searching while specialists shaped prompts, read the output, tried to reproduce it and judged which reports were ready to disclose. It’s that division of labor that makes the effort a human-AI review system instead of a scanner with a press release attached.

$20,000 and 150 repositories
Hamilton said on Aug. 3 that the effort had put over $10,000 into scanning over 100 repositories, and had disclosed critical findings immediately whenever a proof of concept showed exploitability. His Aug. 4 report cited roughly $20,000 in spending, more than a dozen disclosures and 150 repositories scanned.
The scanning kept growing. Outreach, handoff and triage remained described as live operational constraints. Because the snapshots give no comparable disclosure denominator at 55 hours, the rate at which findings landed can’t be set against the rate at which they were resolved.
Hamilton later pointed to the separate Coldcard incident as the catalyst behind the wider campaign. Nothing in the campaign record credits this sprint with discovering the Coldcard flaw.
Most projects have nobody to email
This, I’d argue, is the single most useful thing the sprint turned up. Bitcoin Red Team’s 55-hour update put the share of scanned projects with a SECURITY.md file at 19.5%, and the share with an email inside it at 13.1%.

Left out of the thread were the project corpus, guidance on how to read the denominator, and the measurement method — so those percentages describe this campaign’s scan and nothing wider. Even so: flood a pipeline with findings when the projects at the far end have no documented way to receive them, and volume isn’t what’s holding you back.
The critics have the same evidence problem
According to the developer known as Calle, project owners verified most critical reports quickly. That post came with no denominator, no verified-report count, no rejection count and no patch status, leaving both the breadth and the outcome of that verification unresolved.
From the opposite direction, JW Weatherman argued in public that the campaign was incapable of triaging its own output. He pointed to no campaign-linked issue, patch or advisory, making it criticism with no measurable failure rate behind it.
The same wall stops both claims. With the disposition data unpublished, neither one can be checked.
What a real accounting would look like
Break the findings out by what happened to them: reproduced, acknowledged, downgraded, rejected, fixed. Attach a definition and a denominator to every rate. A breakdown like that would reveal how much of the campaign’s volume converted into actionable work — the line between a security result and a throughput result.
As things stand, 6,700 counts campaign-labeled findings and triage candidates. Speed is what the sprint proved about machine-assisted review. Whether it matters long term rests entirely on the fraction experts can validate, disclose and convert into patches — and that figure is still unpublished.



















STAY ALWAYS UP TO DATE