Use case

How to Refactor Legacy Code with AI Without Changing Behavior

Use AI coding agents to characterize legacy behavior, refactor one seam at a time, verify every diff, and roll out safely without accidental rewrites.

Updated 2026-09-30
AI code refactoringlegacy codecharacterization testssoftware modernization

The safest AI-assisted legacy refactor is deliberately boring: preserve observable behavior, change one boundary, prove equivalence, and keep rollback cheap. Do not ask an agent to “clean up this module.” That gives it permission to reinterpret quirks that callers may depend on.

Use the broader AI coding-agent guide to compare products and the CLI coding-agent shortlist for terminal-first workflows. The process below matters more than the model.

1. Define the behavior that cannot move

Write an explicit preservation contract before editing. Include public function signatures, HTTP status and error shapes, database writes, events, ordering, rounding, serialization, retries, timing-sensitive side effects, and supported inputs. Add non-goals such as “no schema change,” “no dependency upgrade,” and “no performance rewrite.”

Separate desired behavior from observed behavior. Legacy systems often contain bugs that downstream consumers now treat as contracts. Record a suspected bug, but do not silently fix it inside a refactor. Make that a separate change with its own acceptance criteria.

Capture a clean baseline: exact revision, dependency lock, test commands, representative fixtures, and production metrics. The relevant suite must pass on the current implementation. Prove each new characterization test can detect change by deliberately mutating a covered behavior, confirming the test fails, and reverting the mutant. If the relevant baseline remains red, stop rather than asking the agent to refactor through it.

2. Add characterization tests before the refactor

Characterization tests answer “what does this code do now?” Cover happy paths, errors, empty and boundary values, state transitions, and side effects at the module boundary. Golden-master or snapshot tests can help with large structured output, but review volatile fields and mask only values that are truly nondeterministic.

Have the agent generate candidate cases, then make a human verify the fixtures and assertions against current behavior. Commit this test-only change separately. Never let the same refactoring step weaken assertions, refresh snapshots, or rewrite fixtures to make its implementation pass. The detailed unit-test workflow covers the red-green evidence needed here.

Tests written from source alone still miss runtime wiring, dynamic imports, framework magic, and rare traffic. For a high-risk path, capture sanitized real inputs and outputs, then compare both implementations in a pure or isolated read-only harness with fake adapters for writes, charges, messages, and email. Only the serving path may commit side effects. In a Reddit legacy-refactoring discussion, u/AltUniverseHere reported that old-versus-new replay exposed encoding and rare-input cases their authored tests had missed. That is one practitioner report, not a benchmark, but the failure mode is credible and easy to test in your system.

3. Refactor one seam in an isolated branch

Choose a boundary that can be changed and reversed independently: extract one calculation, introduce one adapter, wrap one dependency, or move one parser behind its existing interface. Map its callers and side effects before editing. Keep the old entry point as a compatibility wrapper until direct and shadow verification pass.

Run the agent in a fresh worktree or disposable environment. Deny production credentials, deployment access, schema migration, package publishing, and unnecessary network access. Permission prompts and sandboxes reduce blast radius; they do not prove that a proposed edit preserves semantics.

Use a prompt with an executable contract:

Refactor BillingService.calculate_total by extracting tax calculation.
Preserve every observable result, exception, log/event, database write,
ordering rule, and public signature. Allowed files: billing/service.py and
billing/tax.py. Do not edit tests, fixtures, snapshots, migrations, or APIs.

Before editing: list callers, side effects, ambiguities, and run:
pytest tests/billing/test_service_characterization.py -q

Keep the diff under 200 changed lines. After each logical edit, rerun that
test, then run mypy billing and the billing integration suite. Stop without
editing if baseline tests fail, behavior is ambiguous, another file is needed,
a dependency/API/schema change is required, or the diff exceeds the boundary.
Return the diff summary, exact commands/results, unresolved risks, and rollback.

Treat every stop as useful evidence. Resolve the ambiguity or split the task; do not broaden scope mid-session.

4. Review behavior, not aesthetics

Inspect the complete diff and confirm the agent did not alter the contract tests. Look for changed defaults, exception handling, transaction boundaries, concurrency, query count, security checks, and newly shared mutable state. Run tests independently of the agent’s transcript. Use a second context for review so the reviewer does not inherit the author’s assumptions, then follow the AI pull-request review workflow.

Stop the refactor if verification needs production-only data, the dependency graph is still unknown, the agent repeatedly edits locked files, tests are flaky or too weak to distinguish implementations, or review cannot explain a semantic change. Passing tests are necessary; unexplained behavior is still a blocker.

5. Roll out as a reversible migration

For critical paths, compare old and new implementations side by side while serving the old result. Shadow only pure logic or use an isolated read-only replay with fake side-effect adapters; never send two emails, submit two charges, enqueue twice, or duplicate a write. Log structured, privacy-safe mismatches and investigate them. Then enable the new path for internal traffic or a small canary behind a flag. Watch error rate, latency, resource use, and business invariants. Define the rollback command and owner before increasing exposure. Remove the compatibility path only after a stable observation window and a final dependency search.

Which tool fits this workflow?

ToolBest fitCurrent price and limit reality, checked September 30, 2026
Claude CodeInteractive repository exploration with explicit edit and command approvalsClaude Pro is $20 monthly or $17/month billed annually. Claude Code shares rolling five-hour and weekly plan limits across Claude surfaces; there is no fixed task count. API/provider login is usage-based.
OpenAI CodexIsolated local or cloud tasks with a strong diff and review surfaceChatGPT Plus is $20/month; Pro starts at $100/month. Codex and ChatGPT Work share variable plan usage, with current capacity shown in the dashboard or CLI /status; API-key use is billed separately and excludes cloud features.
AiderTerminal-first, model-flexible work with automatic Git commits and configurable lint/test loopsAider is Apache-2.0 software. Most models require a separate provider key, so provider pricing, rate limits, and context limits apply; Aider does not enforce a universal token cap. Compare its operating model in Aider vs OpenCode.

No tool makes an uncharacterized rewrite safe. The useful differentiator is how cheaply your team can constrain files, preserve tests, inspect diffs, repeat checks, and revert one small step.

Sources & further reading