Use case
How to Generate and Maintain Unit Tests with AI
Use AI coding agents to write durable unit tests with behavior-first acceptance criteria, red-green proof, mutation checks, and controlled dependencies.

Verdict
Use an AI agent to draft a small behavior-based test matrix, then require evidence that its assertions detect regressions. Start from behavior written independently of the implementation. For a bug fix, prove red on the base revision and green after the fix; for already-correct behavior, require a passing baseline and a failing deliberate mutation. Keep a human-owned acceptance artifact and CI outside the agent’s control.
This is a workflow guide, not a claim that one agent writes better tests. See the CLI coding-agent shortlist and broader AI coding-agent guide for product selection.
Choose the control surface
| Tool | Best fit for this workflow | Verified price and limit model |
|---|---|---|
| Claude Code | Interactive repository analysis, test planning, and permission hooks around test edits | Claude Code is included in paid Claude plans. Pro is $17/month billed annually or $20 monthly. Usage resets on a rolling five-hour window, paid plans also have weekly limits, and Claude surfaces share the pool; Anthropic publishes no fixed prompt count. |
| OpenAI Codex | Local or cloud test tasks where diffs, command output, and isolated execution matter | Plus is $20/month and Pro starts at $100/month. Usage varies by model and task; local messages and cloud chats share plan allowance, and weekly limits may apply. Check the live dashboard or CLI /status. |
| Aider | Git-first local pairing with an explicit test command after every edit | Apache-2.0 software with no Aider subscription or quota. You pay the selected model provider, or run a local model. /test runs a command; --test-cmd plus --auto-test can run it after edits. |
For a terminal-only decision, compare Aider and OpenCode. For existing code with unclear seams, start with the AI-assisted legacy refactoring workflow.
1. Write the acceptance artifact before the prompt
A useful test encodes a product decision. Give the agent a small contract that can survive an implementation rewrite:
unit: quoteRenewal
observable_behavior:
- active annual plans receive the renewal discount
- a plan expiring exactly at 2026-10-01T00:00:00Z is not active
- upstream timeout returns retryable_error without charging
invariants:
- amount is never negative
- one request creates at most one charge
out_of_scope: UI copy, database migrations
test_command: npm test -- quoteRenewal.test.ts
Do not begin with “add coverage for quoteRenewal.” That invites tests shaped around branches, private methods, and current mocks. Ask for equivalence classes, boundaries, invariants, and failure paths first. Review that matrix before the agent sees or edits the test file.
2. Require red, then green
For a bug fix, the new regression test must fail on the base revision and pass after the fix. Save the base commit, exact command, exit code, and relevant failure. A syntax error is red but proves nothing; the failure must match the missing behavior.
For new behavior, temporarily revert the production hunk or apply the test to a clean base worktree. If the test still passes, it probably verifies existing behavior, a fixture, or merely that code executes. For coverage of already-correct behavior, a green base is expected. Prove sensitivity with a deliberate behavior-changing mutation instead of manufacturing a base failure.
3. Mutate behavior, not formatting
Use one deliberate mutant before trusting a generated boundary test. In the example above, change the eligibility comparison from expiresAt > now to expiresAt >= now. The exact-boundary test must fail. Revert the mutant and record the result.
Other useful mutations include deleting a validation branch, swapping two status values, removing idempotency lookup, or changing a retryable error to permanent. Surviving mutants reveal weak assertions or an unexecuted path. Run a diff-scoped mutation tool when available; do not claim a mutation score unless the tool actually produced it.
4. Control state, time, randomness, and networks
Inject a clock and freeze it at a named instant. Seed randomness or pass a deterministic ID generator. Give each test a fresh in-memory store, transaction, or temporary directory. Put network calls behind an adapter and use a small fake that returns explicit success, timeout, malformed, and retry responses.
Assert observable output and state transitions: returned value, durable record, emitted domain event, or adapter call at the boundary. Avoid mocking every collaborator and asserting internal call order. Those tests mirror the implementation, fail during harmless refactors, and can remain green when user-visible behavior breaks.
5. Use a prompt that demands evidence
Read the acceptance artifact and current tests. Do not edit production code yet.
Propose the smallest behavior-based test matrix, including boundaries, invalid
inputs, state transitions, time, and network failures. Avoid private-method tests
and mocks of internal collaborators. After approval, add tests only. For a bug fix or new behavior, show the test failing on the base revision for
the expected reason, then passing with the implementation. For existing correct
behavior, record a green baseline. Apply one behavior-changing mutant and show
the relevant test fails,
revert it, and run the targeted suite. Report commands, exit codes, changed files,
and unverified paths. Never weaken an existing assertion to make the suite pass.
The completion artifact should contain: acceptance item → test name, base and changed revisions, red/green output, mutant and kill result, commands with exit codes, files changed, and anything skipped. Review the diff with the AI pull-request review workflow.
6. Maintain tests as behavior changes
Classify failures before editing: product regression, intended contract change, stale fixture, environment failure, or flaky timing. Change an assertion only when the acceptance artifact changed. Delete duplicate and implementation-coupled tests; preserve boundary and invariant coverage. Run fast unit tests on each change, broader integration tests in CI, and mutation checks on changed critical logic or on a schedule.
In a Claude Code discussion, Reddit user u/Aals_aakun described a hook that requires explanation and approval before Claude edits test files, specifically to stop tests being changed to fit wrong code. That is an attributed individual practice, not outcome evidence. Its durable lesson is sound: the actor fixing production code should not be able to move the acceptance boundary silently.