Use case
How to Debug CI Failures with AI Coding Agents
Debug failing CI with AI by preserving the failed run, classifying the failure, reproducing the runner environment, and verifying a minimal fix.

Verdict
Use an AI coding agent to investigate a failed pipeline with its complete evidence. Preserve the failing run and commit, classify the failure, reproduce the environment, and require checks that distinguish competing causes. Accept a fix only when the failure is explained, the relevant check passes, and fresh CI passes on the patched commit.
This workflow starts after a pipeline fails. It is separate from an issue-to-PR task. Compare products in the AI coding-agent guide or CLI coding-agent shortlist.
Choose the control surface
| Tool | Best fit here | Current price and limit model, checked October 2, 2026 |
|---|---|---|
| OpenAI Codex | Reading run evidence with gh, reproducing commands, and reviewing a bounded patch | ChatGPT Plus is $20/month; Pro starts at $100/month. Local and cloud work share variable plan usage, and weekly limits may apply. API-key CI is usage-priced and does not include cloud features. |
| Claude Code | Interactive log investigation with explicit command and edit permissions | Claude Pro is $20 monthly or $17/month with $200 billed annually. Claude Code shares rolling five-hour and weekly limits with other Claude surfaces; Anthropic publishes no fixed message count. API use is billed separately. |
| Aider | Local, Git-first repair when you already know the exact parity command | Apache-2.0 software with no Aider subscription or universal quota. Model-provider charges and limits apply. /test runs a command; --test-cmd plus --auto-test can repeat it after edits. |
1. Freeze the failed run
Record the run URL and ID, attempt, event, head SHA, ref, runner OS and architecture, failing job and step, exit code, and artifacts. Save the complete failed-step log, not a screenshot or stack-trace tail. gh run view RUN_ID --log-failed retrieves failed-step logs; --attempt selects the original attempt.
Check out the SHA in a clean worktree. Read the workflow there because the default branch may already use a different action, runtime, or command. Give the agent secret names and availability rules, never values. Redact tokens, customer data, signed URLs, and environment dumps before adding logs to model context.
2. Classify before changing code
Make the agent assign one primary class and cite the evidence:
| Class | Typical signal | First discriminating check |
|---|---|---|
| Code or test | Same assertion, compiler, lint, or type error repeats | Run the exact step against the failed SHA |
| Dependency or cache | Resolution, checksum, generated-file, or stale-artifact mismatch | Install from the lockfile with an empty cache |
| Environment drift | Only one OS, architecture, runtime, locale, timezone, or service image fails | Print pinned versions and change one differing variable |
| Identity or event context | Fork PR lacks a secret; token, OIDC, or repository permission differs | Inspect event + permission metadata without exposing credentials |
| Nondeterminism or capacity | Failure moves, disappears on rerun, times out, or reports ports, memory, disk, rate limits | Repeat with seed, order, timing, and resource telemetry retained |
| Workflow wiring | Wrong matrix value, conditional, dependency, artifact path, or working directory | Expand the resolved job and trace producer to consumer |
A green rerun does not prove a fix. It changes the leading hypothesis toward flakiness, resource pressure, or an external dependency. GitHub reruns preserve the original GITHUB_SHA and GITHUB_REF, which makes a rerun useful evidence; label it as a new attempt, not a new commit.
3. Reproduce the environment contract
Mirror the job's command, failed SHA, runtime and package-manager versions, lockfile mode, OS/architecture, working directory, environment variable names, services, and relevant locale/timezone. Start with empty dependency and build caches. GitHub-hosted jobs use fresh instances of the selected runner image, while *-latest means GitHub's latest stable image and can move over time; pin versions that affect behavior.
If local reproduction is impossible, run a diagnostic CI commit or manual workflow that changes one observable variable. Enable debug logging only when ordinary logs cannot distinguish setup from execution; extra verbosity can expose sensitive data.
4. Give the agent a falsifiable task
Illustrative prompt; replace the run, revision and environment with your observed values:
Diagnose GitHub Actions run 4812, attempt 1, job test-node-24 at commit abc123.
Evidence: failed-step log at artifacts/4812-test.log; workflow from that commit;
runner ubuntu-24.04; Node 24.6.0; npm ci; command npm test -- billing-date.
Do not edit yet. State the failure class, three ranked hypotheses, evidence for
and against each, and one cheap check that would falsify each. Reproduce the
failure from the clean worktree with the same environment. Then propose the
smallest fix. Do not change tests, lockfiles, action versions, timeouts, retries,
or caches unless the evidence identifies them as the cause. Report every command,
exit code, changed file, unresolved difference, and CI verification needed.
This prevents an “upgrade everything and add a retry” patch that destroys evidence and can leave the defect intact.
5. Work a concrete environment failure
Suppose billing-date.test.ts passes on a developer laptop but CI reports expected 2026-09-30, received 2026-10-01. The input instant is 2026-10-01T00:15:00Z. The agent should inspect the runner's timezone, then run the exact test with TZ=UTC and TZ=America/Los_Angeles. If that flips the result, the cause is an implicit timezone contract, not “flaky CI.”
The minimal repair is to make the business timezone explicit in formatBillingDate(instant, accountTimeZone) and test the boundary under TZ=UTC. Verify the targeted test, the repository's full CI test script in a clean environment, and the patched commit's workflow. Do not solve it by setting the whole runner to a developer's timezone unless that is the product contract.
6. Prove the fix at the original boundary
Keep the patch limited to the diagnosed cause. For deterministic failures, retain the before/after command, exit codes, and relevant output. For flaky failures, preserve a failing seed or schedule, repeat the condition, and report the count without claiming permanent elimination. For cache failures, verify cold and warm paths. For matrix failures, rerun the failed cell and required matrix.
Then inspect the complete diff using the AI pull-request review workflow. If tests changed, apply the separate unit-test evidence workflow. A CI badge proves one run passed; the evidence bundle explains why the repair deserves to merge.
In a Reddit thread, u/aradil reported having Claude Code inspect remote GitHub Actions logs because a locked-down environment could not run mobile tests requiring emulators. That is one practitioner report, not effectiveness evidence. When parity is impossible, keep the agent on authoritative remote logs and name the local gap.
Sources & further reading
- GitHub CLI workflow-run logs
- GitHub Actions debug logging
- GitHub Actions job reruns
- GitHub-hosted runner behavior
- OpenAI Codex non-interactive and CI mode
- OpenAI Codex pricing and usage limits
- Claude pricing and usage limits
- Aider linting and testing documentation
- Aider source repository and license
- Community report about remote CI-only reproduction