Your code hasn't changed, but the tools reviewing it have
Sebastian Banescu·September 30, 2026·
Your code hasn't changed, but the tools reviewing it have
On May 29, 2026, Taylor Hornby found a critical bug in Zcash's Orchard circuit that could have let an attacker mint unlimited, undetectable ZEC inside the Orchard pool. The bug had been live since May 2022. It survived a QEDIT audit that covered the circuit, an NCC Group review of the spec and surrounding code, and review by the cryptographers who designed it.
Taylor found it with an automated audit running on Claude Opus 4.8, which had been released the day before.
In September, Vitalik argued that cybersecurity is "naturally defense-favoring once people get their sh*t together." I agree with him on the direction. But the question every team should be asking after Zcash is: when your review tools improve, who is responsible for revisiting what earlier reviews missed?
What happened
The bug was in halo2's variable-base scalar multiplication gadget. Orchard uses it to check that the keys you spend with actually belong to the address holding the note. The circuit checked that every round of the multiplication used the same base point, but never that it was the right one. Two lines used assign_advice(), which just writes a value, where copy_advice(), which also ties it to the real base point, was needed (advisory, walkthrough).
With the base point free, an attacker could make that check pass with any key. That meant spending the same note over and over, each time with a new nullifier, so every double-spend looked like a normal transaction.
This is the classic zero-knowledge bug, an under-constrained circuit, and it's nasty for two reasons:
- Normal tests pass, because honest inputs satisfy the circuit.
- Reviewers read the constraints that are there. Spotting one that's missing means checking every value against the intended math.
Taylor reported the bug that night with a working proof of concept. Zcash disabled Orchard with an emergency soft fork on June 2 and turned it back on with the fix on June 3 (UTC). Even that didn't go smoothly: the first soft fork failed over a missed peer-banning rule and had to be reissued 60 blocks later.
Zcash's turnstile kept total ZEC supply safe. Honest Orchard holders were less protected, since counterfeit coins cashed out first would leave the pool short. And because Orchard is private by design, nobody can prove the bug was never used. Shielded Labs thinks it's unlikely.
It wasn't just the model
It's tempting to read this as "AI found the bug." Taylor's work log tells a more useful story:
He started building AI auditing tools in September 2025, then a set of agents specifically for Zcash after Shielded Labs hired him. His framework first lists every relevant code location, spec statement, security property and failure mode, then assigns an auditing agent to each item.
The same framework, running on Opus 4.7, had audited Orchard before and missed the bug.
Opus 4.7 did find it when told exactly which gadget to check. With a generic prompt, even Opus 4.8 found it in only 1 of 4 runs. Taylor is clear that a handful of trials isn't science.
One difference that may have mattered: the run that worked was given the halo2 book, not just the protocol spec.
Opus 4.8 was very skeptical of its own finding. It assumed audited code must be correct and needed some prodding to take the bug seriously.
IMO, that's the most useful part of the whole story. The model didn't replace Taylor, it made months of his preparation pay off. It's the same point I made in my article on AI audits: the AI speeds the auditor up, it doesn't replace them.
What would an attacker have needed? Opus 4.8 was public. Taylor's account was on Anthropic's allowlist for security researchers, and he says getting on it took a link to a forum post announcing his hire and I know getting on it requires both passing a KYC as an individual and a KYB for the business that employs you. He also says writing the exploit took little Zcash-specific knowledge, but as an SR using LLMs I see many false-positives on a daily basis often stated in a very confident voice. Therefore, it's unlikely that someone with little Zcash-specific knowledge would have taken a finding that the model was skeptical about seriously.
As far as anyone can tell, defense won this time. It won because someone was already on the job with the right tools when the model shipped, and the ecosystem could ship a fix within days. That's what "getting your act together" looks like in practice: a standing capability, not a one-time audit.
Where formal verification fits
Vitalik's long-run answer is AI-assisted formal verification, and I think that's probably where this lands eventually. Shielded Labs has already started a formal verification project for the Orchard circuit.
But a proof only covers the properties you specify, under the assumptions you state. For Orchard, a proof connecting the implemented constraints to the intended math could have caught this bug. A proof of the spec alone wouldn't have, because the spec was fine.
Most of DeFi isn't close to full verification yet. Many of its critical properties depend on prices, incentives and integrations, which are much harder to write down, and you can't prove a property nobody has written down.
"Do we need to re-audit every time a new model ships?"
I run an audit firm, so "audit more" is a convenient conclusion for me. That's why I'll try to be specific about when another review is actually worth it.
The short answer is no, not everything. But the old model, where you audit once and move on with your life, is dead. Here's how I'd think about it:
Re-check the properties that would sink you on every major model release. For Zcash that's balance integrity. For a lending protocol it might be solvency. For a bridge, that nothing gets released without a matching deposit.
Keep a regular, risk-based review cycle for everything else, and add a review when something changes. The tools got better on code like yours (keep a private set of real bugs from your history and see whether a new model finds what your last setup missed). Your TVL grew well past what it was at audit time. Your code or its dependencies changed, and Orchard's bug sat in a library the circuit imports. Or you can't patch quickly, which puts immutable contracts, circuits and bridges first in line.
Run the review properly. List every property and failure mode before auditing anything, give the model your design docs and not just the spec, and run it more than once. One clean run is a sample, not a verdict. Turn every suspected bug into a working proof of concept. You know the old saying: PoC||GTFO.
Prepare the response before you need it. Who can pause what, who signs, and how fast can a fix ship? Consider a multisig that can switch off a shielded pool, which could cut days of exposure.
Keep in mind that I'm not saying you should re-audit your whole codebase every time a lab ships a model. Targeted, ongoing review is what our Continuous Security engagements are for.
The protocols still standing in five years won't be the ones with the best audit from 2023. They'll be the ones that look again when the tools get better, starting where a bug would hurt the most. Everyone else is waiting to find out whose Zcash moment comes next.
If your last audit predates the current generation of models, or nobody on your team owns the question of what it might have missed, it's probably worth talking to us.
Ship Safely. 🚀