Parallel Claude Code Agents Caught a Bug Our Dry-Run Missed
Parallel Claude Code Agents Caught a Bug Our Dry-Run Missed
TL;DR: We used five Claude Code agents running in parallel to convert a 7-module security course into printable PDF handouts. One of them caught a real bug in the course content — a leaked internal credential path — that a prior manual dry-run had already found and “fixed,” except the fix only landed in one of the two places it needed to.
Tonight was supposed to be a formatting job. The Secure AI Builder Bootcamp — our live cohort course on building and securing AI agents — needed its lesson content turned into PDF handouts, because they post better in the community platform than raw markdown. Twenty-one lessons across five modules. Tedious, mechanical, exactly the kind of work you hand off.
So I handed it off. Five Claude Code agents, one per module, running in parallel, each given the same template and the same instruction: convert every lesson faithfully, and don’t paraphrase commands or code — copy them verbatim, because a student is going to run what’s on the page.
That last instruction is what caught the bug.
What We Were Actually Doing
Each agent’s job was narrow: take the markdown source for its module’s lessons, render it into a branded HTML template (green accent, monospace code blocks, numbered step markers), then print each one to PDF with headless Chrome. Four agents got straightforward conceptual-to-procedural content — secrets management, MCP servers, Terraform. The fifth got Module 6: Build, Attack, and Harden Your AI Agent.
Before any of this, we’d already run a manual dry-run of the whole course — walked it end-to-end as a student would, logged every point of friction in a running gap list. That dry-run had already caught and fixed a real blocker in Module 6: a lesson that told students to fetch their API key from /test-bot/anthropic-key — our own internal AWS SSM namespace, not the student-facing one (/bootcamp/anthropic-key). Classic copy-paste-from-internal-docs leak. It got fixed. The gap list said so. We moved on.
The Bug the Dry-Run Missed
Here’s the thing about “we fixed it”: we fixed it in one lesson. The same exact command — same wrong path, same copy-paste origin — was sitting untouched in a second lesson three files over, in the red-teaming exercise where a student attacks their own agent. Nobody had walked that lesson in the dry-run closely enough to catch it, because the dry-run gap list only documented the first occurrence.
The PDF-conversion agent found it anyway. Not because it was looking for security bugs — because its instructions said “preserve every command verbatim, don’t touch the content, just format it.” When you tell an agent to copy something exactly rather than summarize it, it has to actually read every line closely enough to reproduce it. That’s a different kind of attention than a human skimming a lesson for tone and flow, and it’s a different kind of attention than a human dry-run walking the happy path through a course.
The agent flagged it, unprompted, in its final report: “the same bug is still present, unfixed, in lesson-04-red-teaming.md, line 31… I did not silently correct this — I converted lesson-04 faithfully to match the current (buggy) markdown source… recommend fixing this.” It didn’t fix it itself. It told me exactly where, exactly why, and let me decide. That’s the right call — a formatting task quietly rewriting course security guidance on its own authority would be a worse outcome than a students hitting a AccessDenied error.
I fixed the source markdown, the HTML, both PDFs, and the module bundle, then re-grepped the entire course for other copies of that same leaked path. Clean. Updated the original gap list with the finding so the next person who reads it knows the fix wasn’t complete the first time.
What This Says About Manual Review vs. Repeated Passes
In 25 years of enterprise security work, I’ve watched this exact failure mode play out in code review more times than I can count: a vulnerability gets found, gets fixed in the file someone’s looking at, and the same pattern sits unpatched two files away because nobody re-ran the search after the fix. “Fixed” gets treated as a fact about the codebase instead of a fact about the one place someone looked.
A single manual dry-run is a sample, not a proof. It walks one path through the material and catches what’s on that path. What actually closes the gap is a second independent pass with a different lens — in this case, an agent whose job forced it to read every command character-for-character instead of skimming for meaning. That’s not a replacement for the dry-run; it’s a second, cheap, parallel check that happened to be looking at the content for an unrelated reason and caught something the first pass didn’t.
The practical takeaway isn’t “have AI review everything” — it’s narrower than that: when you fix a leaked credential or a copy-paste bug, grep the whole repo for the literal string before you consider it closed, not just the file where you found it. We should have done that the first time. We’re doing it now as standard practice for the rest of this course.
Key Takeaways
- Parallel Claude Code agents doing unrelated mechanical work (PDF formatting) can surface real bugs as a side effect, when their task requires verbatim accuracy rather than summarization
- A manual dry-run finds what’s on the path it walks — it’s a sample, not a guarantee the same issue isn’t sitting somewhere else in the repo
- After fixing a leaked credential path, grep the entire codebase for that literal string — don’t assume “fixed in the file I found it in” means “fixed”
- Agents that flag a problem and stop, rather than silently fixing content they weren’t asked to touch, are doing the safer thing — even when the fix seems obvious
This is part of our ongoing build-in-public series documenting the Secure AI Builder Bootcamp — a hands-on course teaching AI-first builders to build and secure their own agents on real cloud infrastructure.
Comments