How to reliably fix bugs with Claude Code

In this post, I’ll recommend three ways to fix software bugs with AI, specifically Claude Code:

  1. The quick and riskier way
  2. The human-steered thorough way
  3. The software factory thorough way

This is an upgrade path—you can start at No. 1 and work your way to 3.

The quick and riskier way

Paste the error or bug description into the smartest model you have and tell it “Fix.”

Really.

There is some more nuance to this:

  • You should open Claude Code in the repo of the codebase you are working in, which should have a CLAUDE.md so the agent should efficiently navigate the codebase.
  • It’s helpful to have an MCP for Datadog or your observability platform of choice so Claude can look at logs and related performance
  • It’s also a good idea to have a git workflow skill or instructions for Claude so it knows to apply its fix in a feature branch

My current recommendation is to use Fable on low effort for this.

In my private Greenwald Evals, Fable is just as fast and token efficient as Sonnet low effort and is significantly more capable, passing 14 of 16 evals under almost every condition (no baseline instructions, with a CLAUDE.md base, with a skill), whereas Sonnet low effort passes 10 under specific conditions only.

For low-risk or easily reversible issues (i.e., ones that don’t touch data), given the context I’ve mentioned above, this approach will work just fine. It will at least point you in a helpful direction.

The risk is that Fable is guessing. Because we haven’t asked it to actually reproduce the error or run the code proving the fix, this a theoretical exercise. LLMs can’t actually understand what they are looking at, and everything they produce is theoretical until they see feedback.

That bring us to our next approach.

The thorough way—human-steered

This is a multi-prompt, multi-session approach.

You can use as many bullet points here as you want. Maybe web research isn’t needed, maybe you want to skip the “production-readiness” screens. You can organize this to your taste and token budget.

The essential pieces here are setting a clear goal and proving it with end-to-end feedback against live code:

  • Define the consequence or measurement of the bug
  • Claude must be able to reproduce the bug
  • Claude must prove the fix can’t be broken by running the code
  • The fix delivers a measurable outcome

Discovery

  • Clearly name the problem and its consequence—is what gives Claude its stopping point
  • Reproduce the problem locally
  • Web research best practices for fixing the problem
  • Look across the codebase and other company projects to find the patterns for how this problem is addressed or avoided
  • Propose a solution
  • Build the failing test—capture the bug. The goal of the rest of the session is for the red test to turn green.

Conduct discovery in one session, Claude can (and will) use Explorer agents for the research bullets.

Review the approach

  • Architecture review—is the approach too narrow? LLMs have a tendency to plug a hole with their thumbs.
    • Review the end-to-end path of the issue and all involved consumers
    • Look for siblings and other potential points of failure
    • As a staff engineer, what categorical fix can we make to prevent this problem?
  • Adversarial review—the critical step
    • Claude must build and run the proposed solution and try to break it. Its goal is to find what is wrong.
    • If this fails, the main session must review the issues and see if they are minor fixes that deserve a retry, or if the approach is wrong. Cap review passes at a fixed number, stop the process to avoid an iteration infinity loop on a wrong approach, and then step in to recalibrate (“Take a step back. Based on what we’ve learned so far, what are we missing or should try instead?” Then clear context and take this into a new session.)
    • Do not pass Go, do not collect $200 until the adversarial review passes
    • Be sure to have the adversarial reviewer hand over whatever it built to pick up from and for other agents/sessions to be able to prove

If you get this right, you will fix your bugs.

In this phase, you could have the same Claude session execute these next two steps. It is advisable not to for two reasons:

  • You are going to start running out of context window and Claude will eventually get less intelligent or forgetful, and we haven’t started fixing the problem yet
  • Review and contrary requests should always go in their own sessions. Agents want to do what they are asked—we have already asked Claude for a solution. It will want to make it work. You can improve this behavior with CLAUDE.md but it’s easier to just run separate sessions or sub-agents.
  • Claude sessions can talk to each other. Run these sessions in new windows and have them pass their outputs back to your initial Claude.

Build it

  • Actually build it
  • Test turns green
  • OK, you did it!

This can all happen in your main session.

Production-readiness pass

Review the produced code to taste:

  • style guide/pattern conformance
  • performance
  • security
  • accessibility

Here, you can kick off new sessions to continue the work. Don’t hit clear on your main session until you’ve had the chance to discuss its learnings, in the step below:

Retrospective

  • What could’ve prevented this bug from occurring? How can the agentic harness be improved?
  • Mechanical gates (tests, linters, pre-commit hooks, etc.) are preferable to non-deterministic agent rules

Asked for “improvements,” Claude will provide them, whether they improve things or not. Take care that the recommendations actually solve a problem and not a theoretical case.

You’ll notice by now that this approach largely translates to feature development and not just bugs.

The thorough way: sub-agent workflow

This is what you really want: you put bug in on one end of the pipeline, Claude works through all of this, and comes back in somewhere between 30 minutes to two hours with a solid and tested solution that’s ready for human review without any intervention along the way.

Model choice, effort level, and specific instructions are critical here.

Use Opus 4.8 at high or xhigh effort for orchestration—this is your main working session.

Opus 4.8 is diligent, has high attention to detail and rule-following and really shines in this role. Opus 5 is comparably sloppy and can’t be trusted as a manager and is better at inspired single tasks. Fable is better but not as accurate as Opus 4.8, and more expensive anyway.

What you’ll want to do is wrap this in a slash command (a Claude Code skill) that gives the orchestrator all of the steps and checkpoints that we’ve discussed above. It will be in charge of launching the sub-agents and addressing their output. I use /work as a repo-scoped workflow management command for all of my projects.

Here’s what I have in one of my projects. I have several going with this kind of pipeline and they are all a little different and experimental. See what works for you, tweak and tune until it is successful.

Orchestrator ....................................Opus 4.8 / high
│
├── Discovery (sub-agents)
│   ├── web-research best practices ............ Explorer agent (Opus / high)
│   └── scan codebase for the pattern .......... Explorer agent (Opus / high)
│
├── Review the approach (sub-agents)
│   ├── architecture review .................... Opus / high
│   └── adversarial review (build, run, break) . Opus / high
│
├── Build it (orchestrator)
│   └── red test → green ....................... Opus / high
│
├── Production-readiness pass (sub-agents)
│   ├── style / pattern conformance ........... Sonnet / high
│   ├── performance ........................... Sonnet / high
│   ├── security .............................. Sonnet / high
│   └── accessibility ......................... Sonnet / high
│
└── Retrospective ............................... Opus / high

You can hand over the build step to a Sonnet sub-agent; you can have Fable low give a second opinion about architecture review—whatever Lego blocks you want to try here are probably a good idea and an educational experience for your project.

Let me know what works for you.