The diff that looks like a win

You hand the agent a 40-line helper that's been in production since 2022. You ask it to "clean this up." Thirty seconds later it hands back a 14-line version, well-named, flat, no nested conditionals. The linter is happy. The tests that exist are green.

You merge it. It looks great in the PR.

Six weeks later, a customer on a legacy plan hits a path that hasn't been exercised since the migration they forgot about, and the function returns null where the old code returned a degraded-but-valid response. The original author left a comment that said // old tenants send a string here, don't crash. The agent deleted the comment, the branch, and the test for that branch, because nothing in the current codebase referenced it.

Why this happens

The agent is optimizing for the code as it exists in front of it. It has no idea:

  • which callers are on a version of the client you can't push an upgrade to,
  • which edge case was a hotfix for a specific outage,
  • which input format is still produced by a batch job that runs once a quarter.

It reads the type signatures, the tests, and the current callers. Anything not visible in those three places is invisible to it. So when it simplifies a switch down to the happy path, the dropped arms are gone before anyone asks whether they were still load-bearing.

The cheap guard that catches most of it

Before letting an agent "clean up" a function, I ask for two things in the same diff:

  1. A list of the branches it is deleting, and why each one is safe to drop. If it can't name the caller, the branch stays.
  2. A characterization test for the current behavior, written before the refactor. Pin the input/output pairs of the messy version, then let it refactor against those snapshots. If the snapshots break, the refactor is lying about being a no-op.

The second one is the one I see skipped the most. People assume the existing test suite covers behavior. It usually covers the paths someone remembered to test. The whole reason the function looked messy in the first place is that it grew around cases nobody wanted a test for.

Where I drew the line recently

A similar pattern showed up on the API side: an agent "simplified" a response schema by making a nullable field required, because "every current client sends it." That was true for the three tenants we onboarded in 2025. It was false for the two from 2023 that we'd stopped asking.

We ended up keeping the contract as the source of truth locally and diffing it against the live response before we trusted any refactor — which is the whole reason we built Powerduck the way we did. The point isn't the tool. The point is that a refactor claim has to be checked against something that remembers what the messy code was actually defending.

A question worth asking your reviewer

Next time an agent hands you a prettier function, don't start with "is this easier to read?" Start with: which behaviors did you delete, and how do you know they weren't still load-bearing?

That question is annoying to answer. That's exactly why it works.