I counted Claude Code's mistakes for six months
Between March and September 2026 I built ClearMyInbox, a bulk unsubscribe service, on my own, and Claude Code wrote nearly all of it: 944 merged pull requests. I spot-checked diffs. Automated gates did most of the reviewing. So I went back through the whole history and counted where the agent's mistakes were caught: by CI, by a second Claude reviewing every pull request, by two full-codebase sweeps, or only once the code was in use. This is those numbers, including the ones I would rather not publish.
15 minute read. Drafted by Claude Code from the repository's own history. Every figure was pulled from GitHub, git and SonarQube and checked against source on 19 September 2026. How it was made is at the end.
What I built, and what this post isn't
ClearMyInbox connects to a Gmail, Outlook or IMAP mailbox, finds every newsletter and marketing sender, and unsubscribes from them in bulk. It has OAuth with Google and Microsoft, background scan and unsubscribe jobs, a real browser for the unsubscribe pages that need one, Stripe billing, and the controls Google's CASA assessment asks for. It is my product, so read this knowing that.
I'm a .NET contractor by trade. I chose the architecture (vertical slices, a Result type instead of exceptions, one folder per feature) and it was written into the rules file on day one. The agent wrote the code. This post is about what happens between the agent writing something and it reaching users, counted rather than remembered. It is not a productivity claim: I have no second copy of me who built it by hand, and I say more about that near the end.
The setup: one engineer, one agent, a stack of checks
Claude Code runs in a git worktree per task, reads a CLAUDE.md of repo rules, writes the change and its tests, runs them locally and opens a pull request. From there it has to get through CI (xUnit unit and integration tests, Vitest, Playwright, and every migration applied to a real SQL Server), a SonarQube quality gate, and a second Claude that reviews every pull request against a checklist. CI runs on four self-hosted runners on my laptop, which has its own story.
None of those checks existed at the start. Almost every one arrived because of a specific failure, which is the thread running through the rest of this post:
| Check | Added | Because |
|---|---|---|
| CLAUDE.md: stack and architecture rules | 10 Mar | Day one |
| CI: .NET tests, Playwright, then Vitest | 29 Mar | Pull requests started the next day |
| Whole-codebase sweeps by parallel review agents | 4 and 20 Jul | Getting ready for Google's security review and the launch |
| An AI reviewer and a SonarQube gate on every pull request | 20 Jul | "Every PR merged in this repo to date has done so with zero review passes", as the pull request that added it put it |
| "See the test fail" rule, three hooks, CLAUDE.md rewritten | 17 Aug | Tests that could not fail, and rules restated in chat over and over |
| The reviewer must post a verdict or fail | 15 Sep | It went silent while its check stayed green |
The model changed five times along the way, going by the co-author lines in the commits: Opus 4.6 in March, Opus 4.7 in April and May, Opus 4.8 in June and July, Fable 5 and Opus 5 through July and August, and Fable 5.1 in September. The reviewer runs Opus 4.8.
The numbers
332 pull requests were fixes and 200 were features. The rest were docs, CI, infrastructure, refactors and tests, plus 175 titled in plain English with no prefix. The median one added 190 lines, removed 11 and touched 4 files. The median time from opening to merging was 25 minutes, and 59 merged on the busiest day. In total they added 485,216 lines and removed 70,434.
That is an activity number, not an output number, and I would distrust any post that led with it. Pull requests here are small on purpose, and a lot of those lines are tests. The codebase today is 58,431 lines of C# in the API against 103,062 lines of C# tests, plus 56,327 lines of TypeScript in the web app. There are 3,959 .NET test methods, 2,950 Vitest cases and 448 Playwright tests. SonarQube reports 93.5% coverage on the API and 88.9% on the web app, with no open bugs or vulnerabilities on either.
The timestamps also answer the question of who reviewed what. 108 pull requests merged between midnight and 6am, and 35% merged at weekends. I spot-checked the risky ones. The rest went through on the gates.
The sweeps: 117 issues, and the same bugs twice
Before any per-PR review existed, I ran two whole-codebase reviews. On 4 July an audit filed 32 issues. On 20 July, three days after Google's assessment letter, ten review agents ran in parallel across auth, OAuth, the scan pipeline, unsubscribe execution, billing, the data layer, the frontend, infrastructure, test coverage and repo hygiene. They filed 85 issues: 14 critical, 13 high, 46 medium and 12 low. Two agents retracted or corrected their own first findings when asked to re-read the code before quoting it.
The critical ones were live in production:
- Money. An unpaid or abandoned Stripe checkout was recorded as an active paid subscription, and a lifetime purchase the same way, because the payment status was never read. A retried webhook could bring a cancelled subscription back to life. Two copies of the same webhook arriving together could refund a customer twice.
- Users' data. Connecting a second Gmail account silently overwrote the first. A temporary 429 from Google's token endpoint deleted the user's connection permanently. A stuck scan lock could stop a user ever scanning again.
- Infrastructure. Every production secret sat in cleartext in the Terraform state blob, in a storage account that still allowed shared-key access: the JWT signing key, the Stripe secret and webhook keys, the OAuth token encryption key and the rest.
- The agent-era one. When an unsubscribe link can't be found directly, ClearMyInbox asks a model to find it in the email's HTML. Nothing checked that the link it returned was really in the email, so a crafted email could put an attacker's link in the dashboard as a trusted unsubscribe link.
- The safety net. Stripe's webhook signature check was replaced by a stub in every test, so none of the billing tests exercised it.
84 were fixed and one was closed as not needed, all within two days. The part that changed how I work is below. Four of the July problems had already been reported on 4 July:
| Reported on 4 July | Then | Found on 20 July |
|---|---|---|
invoice.paid can resurrect a cancelled subscription | Closed as fixed, 5 Jul | Two other webhook handlers can resurrect a cancelled subscription. The first one had the guard; its siblings did not. |
| A leaked scan lock locks a user out until restart | Closed as fixed, 5 Jul | An orphaned scan lock blocks a user permanently |
| One Gmail token is reused for all Gmail connections | Closed as fixed, 5 Jul | A second Gmail connection overwrites the first |
| Key Vault secrets inlined into job specs and Terraform state | Half fixed, and a code comment said it was done | Every production secret still in cleartext in Terraform state. The comment saying otherwise was false. |
Where a fix had shipped, it did what its ticket said. The agent fixed the instance it was pointed at, the ticket closed, and the same mistake stayed alive next door. Nobody had asked it to look for the class. The Terraform one went further: the fix left a comment in the module saying the state no longer held the secret set, which was not true, and a comment like that is what stops the next reviewer looking.
Then speed made it worse. The fix sprint merged 59 pull requests on 20 July and 51 on the 21st. A four-agent re-scan of just those fixes found 10 new problems, most of them introduced by the fixes. One was a health check that opened a database connection every time the uptime monitor pinged, which stopped the serverless database from ever pausing. That was the exact signature of a June incident that had cost about £100 of a £166 monthly bill. Another meant a failed refund could be recorded as a successful one. A third let re-running an older CI build roll production backwards.
The sweep was not always right either. One of its 14 criticals said the production database was reachable from the internet with the SQL admin login enabled. That had been fixed and deployed at 07:40 that morning, and the agent was reading a stale checkout: the review's own notes recorded that the working copy was 17 commits behind. Checking the live server the next day took minutes, and nothing needed changing. The public endpoint is still reachable from any Azure tenant, which is a deliberate trade-off rather than an oversight: private networking costs money this project does not spend, so the database is defended by Entra-only authentication, Defender for SQL and auditing instead.
The reviewer on every pull request
The same day the second sweep reported, I added a review workflow: a second Claude reads each pull request's diff against a checklist and leaves inline comments and a verdict. Its instructions do most of the work:
# .github/workflows/claude-review.yml (the reviewer's instructions, trimmed)- Lead with correctness, security, and data-isolation issues. A missing user-id filter on a query is a blocking bug, not a nitpick.- Every issue you raise must name a concrete failure: the input, state, or sequence that makes it go wrong. If you cannot describe how it breaks, do not raise it.- Do not comment on formatting, import order, or anything a linter owns.- If the change is clean, say so briefly in the posted summary. A short review is a good outcome. Do not manufacture findings to look thorough.Of the 684 pull requests merged from then until 19 September, 199 got at least one inline finding: 353 in all, 337 after removing repeats from re-reviews. 57 verdicts said the pull request should not merge as it was. 92% of the findings were fixed. 22 were declined, and every declined blocking finding was the reviewer being wrong. In total about 2% were wrong on substance. It claimed a build would break (the reply ran the build, which passed), and twice it insisted a React type did not exist when it did.
The findings I value most are the ones no test in the repo could have caught:
- An SSRF hidden by its own comment (#717). A new in-page request followed redirects, which walked straight past the guard against requests to internal addresses. The guard's comment said it closed exactly that hole. The reply: "Correct, and exploitable as described".
- A quota bypass (#973). A fix to refund scans that never finished also refunded scans that had finished, if the user closed the tab at the right moment. It was a way round the free tier's scan limit.
- A silent gap in IMAP scans (#1435). A resumed scan used an email's
Dateheader as its bookmark, but IMAP searches by delivery date. Mail could be skipped while the window was reported as fully scanned. - A CI step that did nothing (#943). A new deploy-tracking call built its JSON wrongly and got a 422 that was swallowed, so the mechanism never engaged while every job stayed green.
- A race the tests couldn't see (#799). Four parallel lookups shared one database context, which throws under real concurrency. The unit tests used mocks that are thread-safe, so they passed.
By a rough keyword count, about a fifth of the findings pointed at a comment, doc or PR description that the code contradicted, and nearly as many argued the existing tests would not catch the problem. Both are exactly what an agent produces fluently: confident prose, and tests that agree with the code they were written beside.
Two layers did less than I expected. CI failed 105 runs on the agent's branches, out of 1,443 that finished, and some of those were runner faults rather than code. SonarQube's pull request gate failed once in 1,396 analyses. It earned its place on main instead, where a September clean-up took 56 open API findings to zero. CI failures on main itself fell from 6.7% of finished runs before 20 July to 2.7% after, though too much else changed at the same time to credit any single check.
What got past everything
Between July and September, 60 more issues were labelled as bugs, apart from the sweeps and the re-scan. Five were flaky tests. Some were subtle. Many were the kind no gate was looking for, because each check only asked whether the code did what the pull request said:
- A Change Password button on the account page that did nothing. The feature behind it had never been built.
- A daily background job that failed every day for 13 days, because a dependency was only registered when the app ran as the API.
- A conversion upload that reported a green heartbeat for two days while uploading nothing.
- One over-long email header that made the database reject a row and threw away a user's entire scan.
- Placeholder "Demo addresses" copy from a test fixture, shown to real customers on the connect screen.
- Google Workspace and Microsoft 365 users on their own domains sent down the manual IMAP path instead of OAuth.
None of them failed a check on the way in. Several were things reporting success: a job, a heartbeat, a scan. I have since learned to treat "green" from anything the agent built as a claim to check, which is also the next section.
When the checks themselves failed
The silent reviewer. On 15 September the review check went green twice without posting anything. In its automation mode, the review action throws away the model's final message. My prompt said "if the change is clean, say so briefly and stop", so on clean commits the verdict went nowhere, and the check passed either way. 60 of the pull requests merged since 20 July have no review output at all, and I can't now tell which of those were reviewed silently and which weren't reviewed. The fix (#1647) makes the reviewer post its verdict with gh pr comment, and adds a step that fails the job when no comment appears. A lost review now shows up red.
Tests that could not fail. The 4 July audit found end-to-end specs in a folder the test runner never looked in. The July sweep found the Stripe signature check stubbed out in every test. The reviewer kept finding tests written alongside the code that would stay green with the fix removed. On 17 August this went into the rules file:
## Evidence (CLAUDE.md, since 17 August)- Never claim a test guards something until you have SEEN it fail without the fix. Every new test, not just ones labelled "guard".- Prove the mutation landed on disk before trusting a green result.- A green path-filtered job that logged "Nothing to test" is not evidence.- Restore mutated files with git checkout --, never mv file.bak file - the stale mtime makes dotnet test re-run the mutated assembly.- Commit before mutating, so restoring does not revert the real work.It still earns its keep. In the pull request for this site's previous post, the first version of a new test passed with the code it guarded deleted. The live data never reached the branch it was meant to test. The rule made the agent delete the guard, watch nothing fail, and rewrite the test until it did.
Mutation testing can bite back. Breaking code on purpose to prove a test works is itself something an agent can get wrong. In one pull request (#1003), an unanchored sed during a mutation run flipped two lines instead of one and was committed. In another (#1465), restoring after a mutation reverted the real fix before it was committed, and the pull request was about to merge a claimed fix with no fix in it. The reviewer caught both. Neither was a bug in the product. Both were bugs in the agent's process, and that is the category most of my rules now cover.
What I still did
I decided what to build and in what order, set the architecture, ran the sweeps, and chose what the gates enforce. I spot-checked diffs, mostly the risky ones, and said no to 16 pull requests, which closed unmerged.
By August, CLAUDE.md was 492 lines, mostly describing the codebase. The rewrite on 17 August cut it to 239 lines of what the pull request called "the rules we kept having to restate". Its summary was blunt: every issue that had recurred over two months was a process failure. Examples were branching from a stale local copy, calling a task done when its pull request went green but before it deployed, and mutation tests that never actually landed on disk. The detail moved into four files under docs/rules/ and six skills that load only when a task needs them.
The same change added three hooks, because, in that pull request's words, "Prose rules get skipped under pressure." One refuses scratch files in the repo root, which then held seven of them. One refuses jq, which isn't installed on this machine and fails silently inside loops. One refuses PowerShell writes into source folders, which had silently corrupted four files with a byte-order mark. Each is a small Python script, and each replaced a rule that had been written down and ignored.
What I can't claim
- That it was faster. I have no control group. METR's 2025 study of experienced open-source developers found they took 19% longer with AI tools while believing they had been faster. I have no measurement that says my case was different, so I'm not claiming one.
- That it will last. Six months is too short to know whether this codebase ages well. SonarQube is clean today, and I have not read every module end to end.
- That it generalises. This is one person, one product and a model that changed five times.
- That the security assessment covers the code. CASA tests controls and configuration. The sweep that found 14 critical problems ran three days after the assessment letter.
- What it cost. I've left cost out of this post.
If you're doing this alone
- Write a rule every time the agent repeats a mistake, and make it a hook or a check once the written rule has been ignored.
- When you file a bug, ask where else it lives. Three of the problems my second sweep found had been closed as fixed a fortnight earlier.
- Make every test fail once before you trust it, and let the agent prove it rather than tell you.
- Verify the verifier. A review, a heartbeat or a job that can go quiet while showing green eventually will.
- Slow down after a big fix batch. My two fastest days produced 10 new bugs, found only because I re-scanned the fixes.
- Count what reached users, not what merged.