Skip to content
By Chris Devine · Updated

I counted Claude Code's mistakes for six months

Between March and September 2026 I built ClearMyInbox, a bulk unsubscribe service, on my own, and Claude Code wrote nearly all of it: 944 merged pull requests. I spot-checked diffs. Automated gates did most of the reviewing. So I went back through the whole history and counted where the agent's mistakes were caught: by CI, by a second Claude reviewing every pull request, by two full-codebase sweeps, or only once the code was in use. This is those numbers, including the ones I would rather not publish.

15 minute read. Drafted by Claude Code from the repository's own history. Every figure was pulled from GitHub, git and SonarQube and checked against source on 19 September 2026. How it was made is at the end.

What I built, and what this post isn't

ClearMyInbox connects to a Gmail, Outlook or IMAP mailbox, finds every newsletter and marketing sender, and unsubscribes from them in bulk. It has OAuth with Google and Microsoft, background scan and unsubscribe jobs, a real browser for the unsubscribe pages that need one, Stripe billing, and the controls Google's CASA assessment asks for. It is my product, so read this knowing that.

I'm a .NET contractor by trade. I chose the architecture (vertical slices, a Result type instead of exceptions, one folder per feature) and it was written into the rules file on day one. The agent wrote the code. This post is about what happens between the agent writing something and it reaching users, counted rather than remembered. It is not a productivity claim: I have no second copy of me who built it by hand, and I say more about that near the end.

The setup: one engineer, one agent, a stack of checks

Claude Code runs in a git worktree per task, reads a CLAUDE.md of repo rules, writes the change and its tests, runs them locally and opens a pull request. From there it has to get through CI (xUnit unit and integration tests, Vitest, Playwright, and every migration applied to a real SQL Server), a SonarQube quality gate, and a second Claude that reviews every pull request against a checklist. CI runs on four self-hosted runners on my laptop, which has its own story.

The layers between Claude Code and production. CI failed 105 runs on the agent's pull request branches. The SonarQube gate failed once in 1,396 analyses. The AI reviewer left 353 inline findings and 57 blocking verdicts. I spot-checked and merged, and closed 16 pull requests unmerged. Codebase sweeps filed 32, 85 and 10 issues. 60 more bugs turned up in use.
Each layer reports in its own unit (failed runs, findings, issues), so compare across a row, not down the page.

None of those checks existed at the start. Almost every one arrived because of a specific failure, which is the thread running through the rest of this post:

CheckAddedBecause
CLAUDE.md: stack and architecture rules10 MarDay one
CI: .NET tests, Playwright, then Vitest29 MarPull requests started the next day
Whole-codebase sweeps by parallel review agents4 and 20 JulGetting ready for Google's security review and the launch
An AI reviewer and a SonarQube gate on every pull request20 Jul"Every PR merged in this repo to date has done so with zero review passes", as the pull request that added it put it
"See the test fail" rule, three hooks, CLAUDE.md rewritten17 AugTests that could not fail, and rules restated in chat over and over
The reviewer must post a verdict or fail15 SepIt went silent while its check stayed green

The model changed five times along the way, going by the co-author lines in the commits: Opus 4.6 in March, Opus 4.7 in April and May, Opus 4.8 in June and July, Fable 5 and Opus 5 through July and August, and Fable 5.1 in September. The reviewer runs Opus 4.8.

The numbers

Bar chart of pull requests merged per week from 30 March to 19 September 2026. Weekly counts stay under 40 except one week of 76 in late April, until the week of 20 July, which has 148. Volume stays around 100 a week until mid-August and then drops.
580 of the 944 pull requests merged in the five weeks from 20 July, the day the second sweep reported and per-PR review was switched on.

332 pull requests were fixes and 200 were features. The rest were docs, CI, infrastructure, refactors and tests, plus 175 titled in plain English with no prefix. The median one added 190 lines, removed 11 and touched 4 files. The median time from opening to merging was 25 minutes, and 59 merged on the busiest day. In total they added 485,216 lines and removed 70,434.

That is an activity number, not an output number, and I would distrust any post that led with it. Pull requests here are small on purpose, and a lot of those lines are tests. The codebase today is 58,431 lines of C# in the API against 103,062 lines of C# tests, plus 56,327 lines of TypeScript in the web app. There are 3,959 .NET test methods, 2,950 Vitest cases and 448 Playwright tests. SonarQube reports 93.5% coverage on the API and 88.9% on the web app, with no open bugs or vulnerabilities on either.

The timestamps also answer the question of who reviewed what. 108 pull requests merged between midnight and 6am, and 35% merged at weekends. I spot-checked the risky ones. The rest went through on the gates.

The sweeps: 117 issues, and the same bugs twice

Before any per-PR review existed, I ran two whole-codebase reviews. On 4 July an audit filed 32 issues. On 20 July, three days after Google's assessment letter, ten review agents ran in parallel across auth, OAuth, the scan pipeline, unsubscribe execution, billing, the data layer, the frontend, infrastructure, test coverage and repo hygiene. They filed 85 issues: 14 critical, 13 high, 46 medium and 12 low. Two agents retracted or corrected their own first findings when asked to re-read the code before quoting it.

The critical ones were live in production:

84 were fixed and one was closed as not needed, all within two days. The part that changed how I work is below. Four of the July problems had already been reported on 4 July:

Reported on 4 JulyThenFound on 20 July
invoice.paid can resurrect a cancelled subscriptionClosed as fixed, 5 JulTwo other webhook handlers can resurrect a cancelled subscription. The first one had the guard; its siblings did not.
A leaked scan lock locks a user out until restartClosed as fixed, 5 JulAn orphaned scan lock blocks a user permanently
One Gmail token is reused for all Gmail connectionsClosed as fixed, 5 JulA second Gmail connection overwrites the first
Key Vault secrets inlined into job specs and Terraform stateHalf fixed, and a code comment said it was doneEvery production secret still in cleartext in Terraform state. The comment saying otherwise was false.

Where a fix had shipped, it did what its ticket said. The agent fixed the instance it was pointed at, the ticket closed, and the same mistake stayed alive next door. Nobody had asked it to look for the class. The Terraform one went further: the fix left a comment in the module saying the state no longer held the secret set, which was not true, and a comment like that is what stops the next reviewer looking.

Then speed made it worse. The fix sprint merged 59 pull requests on 20 July and 51 on the 21st. A four-agent re-scan of just those fixes found 10 new problems, most of them introduced by the fixes. One was a health check that opened a database connection every time the uptime monitor pinged, which stopped the serverless database from ever pausing. That was the exact signature of a June incident that had cost about £100 of a £166 monthly bill. Another meant a failed refund could be recorded as a successful one. A third let re-running an older CI build roll production backwards.

The sweep was not always right either. One of its 14 criticals said the production database was reachable from the internet with the SQL admin login enabled. That had been fixed and deployed at 07:40 that morning, and the agent was reading a stale checkout: the review's own notes recorded that the working copy was 17 commits behind. Checking the live server the next day took minutes, and nothing needed changing. The public endpoint is still reachable from any Azure tenant, which is a deliberate trade-off rather than an oversight: private networking costs money this project does not spend, so the database is defended by Entra-only authentication, Defender for SQL and auditing instead.

The reviewer on every pull request

The same day the second sweep reported, I added a review workflow: a second Claude reads each pull request's diff against a checklist and leaves inline comments and a verdict. Its instructions do most of the work:

# .github/workflows/claude-review.yml (the reviewer's instructions, trimmed)- Lead with correctness, security, and data-isolation issues. A missing  user-id filter on a query is a blocking bug, not a nitpick.- Every issue you raise must name a concrete failure: the input, state,  or sequence that makes it go wrong. If you cannot describe how it  breaks, do not raise it.- Do not comment on formatting, import order, or anything a linter owns.- If the change is clean, say so briefly in the posted summary. A short  review is a good outcome. Do not manufacture findings to look thorough.

Of the 684 pull requests merged from then until 19 September, 199 got at least one inline finding: 353 in all, 337 after removing repeats from re-reviews. 57 verdicts said the pull request should not merge as it was. 92% of the findings were fixed. 22 were declined, and every declined blocking finding was the reviewer being wrong. In total about 2% were wrong on substance. It claimed a build would break (the reply ran the build, which passed), and twice it insisted a React type did not exist when it did.

Horizontal bar chart of 353 AI review findings by category: correctness bug 171, comment or doc contradicting the code 36, CI and infrastructure 32, test would not catch it 25, error handling 20, security or privacy 15, performance 13, concurrency 11, style 11, other 10, API or compatibility 7, data loss 2. Almost every bar is mostly fixed.
Half the findings were plain correctness bugs. The second-largest group was code whose own comments or docs said something untrue.

The findings I value most are the ones no test in the repo could have caught:

By a rough keyword count, about a fifth of the findings pointed at a comment, doc or PR description that the code contradicted, and nearly as many argued the existing tests would not catch the problem. Both are exactly what an agent produces fluently: confident prose, and tests that agree with the code they were written beside.

Two layers did less than I expected. CI failed 105 runs on the agent's branches, out of 1,443 that finished, and some of those were runner faults rather than code. SonarQube's pull request gate failed once in 1,396 analyses. It earned its place on main instead, where a September clean-up took 56 open API findings to zero. CI failures on main itself fell from 6.7% of finished runs before 20 July to 2.7% after, though too much else changed at the same time to credit any single check.

What got past everything

Between July and September, 60 more issues were labelled as bugs, apart from the sweeps and the re-scan. Five were flaky tests. Some were subtle. Many were the kind no gate was looking for, because each check only asked whether the code did what the pull request said:

None of them failed a check on the way in. Several were things reporting success: a job, a heartbeat, a scan. I have since learned to treat "green" from anything the agent built as a claim to check, which is also the next section.

When the checks themselves failed

The silent reviewer. On 15 September the review check went green twice without posting anything. In its automation mode, the review action throws away the model's final message. My prompt said "if the change is clean, say so briefly and stop", so on clean commits the verdict went nowhere, and the check passed either way. 60 of the pull requests merged since 20 July have no review output at all, and I can't now tell which of those were reviewed silently and which weren't reviewed. The fix (#1647) makes the reviewer post its verdict with gh pr comment, and adds a step that fails the job when no comment appears. A lost review now shows up red.

Tests that could not fail. The 4 July audit found end-to-end specs in a folder the test runner never looked in. The July sweep found the Stripe signature check stubbed out in every test. The reviewer kept finding tests written alongside the code that would stay green with the fix removed. On 17 August this went into the rules file:

## Evidence   (CLAUDE.md, since 17 August)- Never claim a test guards something until you have SEEN it fail without the  fix. Every new test, not just ones labelled "guard".- Prove the mutation landed on disk before trusting a green result.- A green path-filtered job that logged "Nothing to test" is not evidence.- Restore mutated files with git checkout --, never mv file.bak file - the  stale mtime makes dotnet test re-run the mutated assembly.- Commit before mutating, so restoring does not revert the real work.

It still earns its keep. In the pull request for this site's previous post, the first version of a new test passed with the code it guarded deleted. The live data never reached the branch it was meant to test. The rule made the agent delete the guard, watch nothing fail, and rewrite the test until it did.

Mutation testing can bite back. Breaking code on purpose to prove a test works is itself something an agent can get wrong. In one pull request (#1003), an unanchored sed during a mutation run flipped two lines instead of one and was committed. In another (#1465), restoring after a mutation reverted the real fix before it was committed, and the pull request was about to merge a claimed fix with no fix in it. The reviewer caught both. Neither was a bug in the product. Both were bugs in the agent's process, and that is the category most of my rules now cover.

What I still did

I decided what to build and in what order, set the architecture, ran the sweeps, and chose what the gates enforce. I spot-checked diffs, mostly the risky ones, and said no to 16 pull requests, which closed unmerged.

Step chart of the length of CLAUDE.md: 291 lines on 10 March, rising to 478 in May and 492 by mid-August, then cut to 239 on 17 August and 252 a day later.
The rules file got shorter once it stopped describing the code and started describing the work.

By August, CLAUDE.md was 492 lines, mostly describing the codebase. The rewrite on 17 August cut it to 239 lines of what the pull request called "the rules we kept having to restate". Its summary was blunt: every issue that had recurred over two months was a process failure. Examples were branching from a stale local copy, calling a task done when its pull request went green but before it deployed, and mutation tests that never actually landed on disk. The detail moved into four files under docs/rules/ and six skills that load only when a task needs them.

The same change added three hooks, because, in that pull request's words, "Prose rules get skipped under pressure." One refuses scratch files in the repo root, which then held seven of them. One refuses jq, which isn't installed on this machine and fails silently inside loops. One refuses PowerShell writes into source folders, which had silently corrupted four files with a byte-order mark. Each is a small Python script, and each replaced a rule that had been written down and ignored.

What I can't claim

If you're doing this alone