Bitcoin’s AI security sprint found 6,700 issues in 55 hours, but no one knows how many are real

Liam 'Akiba' Wright


AI-assisted security campaign focused on the Bitcoin ecosystem, Bitcoin Red Team, said it generated 6,700 findings across 425 projects in its first 55 hours. The campaign labeled 1,029 of them high or critical.

The Aug. 6 update measures how much material entered a security triage pipeline, and its effect on software security remains unreported.

The retrieved thread omitted audit-ready definitions and denominators for the severity counts, as well as case-level outcomes, an aggregate false-positive rate, and a fix rate.

Those missing fields prevent a calculation of how many alerts became confirmed vulnerabilities, how many maintainers rejected or downgraded, and how many led to patches.

Tokenmetrics

The first 55 hours still reveal a consequential capability, noting how AI systems can fill an ecosystem-scale review pipeline quickly. Expert prompting, reproduction, disclosure, and maintainer response remained necessary at every later stage.

What the campaign numbers measure

The campaign published two snapshots as its roster and workload expanded:

Elapsed timeProjectsTotal findingsReported severityParticipants27.5 hours3904,96285 critical; 635 high1655 hours4256,7001,029 high or critical24 reported, including three bots

The 27.5-hour update covered 390 projects and 4,962 findings. By the 55-hour mark, the project count had risen by 35 and the finding count by 1,738. The later thread put high-or-critical findings at 15.4% of the total and clarified that three of the 24 reported participants were bots.

The earlier post separated critical and high findings, while the later one combined them, with both sets of figures reflecting campaign assessments. Maintainer-confirmed exploitability and remediation outcomes require separate evidence.

Infographic showing Bitcoin Red Team's campaign-reported 55-hour totals, human review workflow, disclosure contact gaps, and unpublished false-positive and fix rates.
Bitcoin Red Team scanned 425 projects and reported 6,700 findings, including 1,029 high or critical issues, while public validation rates remain unpublished.

Rob Hamilton described Kimi K3 as handling the heavy analysis, with GPT Sol, Fable/Opus, and GLM 5.2 supporting the documentation. He said OpenAI’s Cyber Harness covered selected components he considered load-bearing.

A day later, Hamilton wrote that subject-matter experts could change an assessment with one or two sentences of context or a small block of code. In examples he described, that input pushed middling concerns into high or critical territory. He also identified operations, disclosure handoff, and triage as bottlenecks.

In Hamilton’s account, models searched broadly while specialists shaped prompts, interpreted output, attempted reproduction, and decided which reports were ready for disclosure. That division of labor makes the campaign a human-AI review system.

OpenAI’s new cybersecurity push has a lesson for crypto: stop waiting for the hackOpenAI’s new cybersecurity push has a lesson for crypto: stop waiting for the hack
Related Reading

OpenAI’s new cybersecurity push has a lesson for crypto: stop waiting for the hack

OpenAI’s Daybreak may point to the crypto industry’s next security standard of becoming resilient before vulnerabilities are exploited.

May 12, 2026 · Gino Matos

The developer known as Calle said most critical reports were quickly verified by project owners. The post supplied no denominator, verified-report count, rejection count, or patch status, leaving the breadth and outcome of that verification unresolved.

CryptoSlate Daily Brief

Daily signals, zero noise.

Market-moving headlines and context delivered every morning in one tight read.

5-minute digest 100k+ readers

Free. No spam. Unsubscribe any time.

Whoops, looks like there was a problem. Please try again.

You’re subscribed. Welcome aboard.

Firefox finds 20 year old bug and patches 14 months of fixes in 30 days using Anthropic’s Mythos AIFirefox finds 20 year old bug and patches 14 months of fixes in 30 days using Anthropic’s Mythos AI
Related Reading

Firefox finds 20 year old bug and patches 14 months of fixes in 30 days using Anthropic’s Mythos AI

Mozilla’s 20-year Firefox bug shows the risk of AI-accelerated zero-day discovery

May 10, 2026 · Liam ‘Akiba’ Wright

Outreach and outcomes define the security value

In the 55-hour update, Bitcoin Red Team reported that 19.5% of scanned projects had a SECURITY.md file and 13.1% had an email there. The retrieved thread omitted the project corpus, denominator interpretation, and measurement method, so the percentages only describe the campaign’s scan.

On Aug. 3, Hamilton said the effort had spent over $10,000 scanning over 100 repositories and had immediately disclosed critical findings when a proof of concept demonstrated exploitability. On Aug. 4, he reported about $20,000 in spending, more than a dozen disclosures and 150 repositories scanned.

Scanning continued to expand, while the campaign described outreach, handoff and triage as active operational constraints. The published snapshots offer no comparable disclosure denominator at 55 hours, so they cannot establish the relative speed of scanning and resolution.

Hamilton later identified the separate Coldcard incident as a catalyst for the wider campaign. The campaign record attributes no discovery of the Coldcard flaw to this sprint.

Coldcard’s $89M wallet bug triggers the biggest Bitcoin movement since FTX and completely distorts market signalsColdcard’s $89M wallet bug triggers the biggest Bitcoin movement since FTX and completely distorts market signals
Related Reading

Coldcard’s $89M wallet bug triggers the biggest Bitcoin movement since FTX and completely distorts market signals

More than 77,000 BTC moved from older wallets as users raced to secure funds, complicating bearish readings across key on-chain indicators.

Aug 2, 2026 · Oluwapelumi Adejumo

A useful public accounting would separate findings that were reproduced, acknowledged, downgraded, rejected, and fixed, with definitions and denominators for each rate. That breakdown would show how much of the campaign’s volume became actionable security work.

A public critic, JW Weatherman, argued that the campaign could not triage its output. His post identified no campaign-linked issue, patch, or advisory, so it supplies criticism without a measurable failure rate. The campaign’s missing disposition data leaves the underlying question open.

For now, 6,700 represents campaign-labeled findings and triage candidates. The sprint demonstrated the speed of machine-assisted review. Its lasting security value depends on the share that experts can validate, disclose, and convert into fixes.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *

Pin It on Pinterest