The 5-prompt debug loop
Five copy-ready prompts that make a coding agent explain, narrow, test and prove a bug fix instead of guessing.
“Fix it” is not an instruction. It is a wish.
Without the full error, the history and a way to test each theory, a coding agent has to guess. Every guess can touch more code, and every extra change makes the original failure harder to see. That is how one broken thing becomes five.
This five-prompt loop replaces the wish with an investigation. Keep the loop nearby, then use the sections below to grade the answer before you move to the next prompt.
The loop card
1. Understand
Here is the full error, including the stack trace:
[paste the complete error]
Right before it appeared, I did:
[what you did]
Do not change code yet. Explain what this error means in plain language. Name the
specific thing that failed and the layer it came from: browser, server, database,
build, infrastructure or a dependency. Tell me what evidence supports that answer.
2. Narrow
Find what changed between the last known working state and this failure.
Inspect the recent commits, the current diff, configuration changes and relevant
runtime or dependency changes. List the likely changed inputs newest first, with one
line explaining why each could affect this error.
Do not change code.
3. Rank
Give me the three most likely causes, most likely first.
For each cause, show:
1. the file, line, setting or runtime input you suspect,
2. the evidence that points to it,
3. one read-only check that can confirm or eliminate it.
Do not guess to fill the list. If the evidence supports fewer than three causes, say
so and tell me what information is missing.
4. Fix one cause
Use the evidence to address cause number one only.
Propose the smallest coherent fix. Show me the diff and explain what every changed
file contributes before applying it. Do not include unrelated cleanup. List anything
else you notice separately and leave it unchanged.
5. Prove
Run the checks that can prove the original failure is gone and that nearby behavior
still works. Show me each command and its result.
If there are no automated tests, give me exact manual steps with the expected result
at every step. Include one negative case that should still fail, so the check cannot
pass regardless of behavior. Report anything you could not verify.
Still broken? Preserve the new evidence and start again at prompt one. After three laps with no progress, stop looping and switch strategy.
Before you start, capture a real error
The loop is only as useful as the evidence you give it. One red line is rarely enough.
Copy the complete message and stack trace, plus the nearby output that appeared just before it. Record the exact action that produced it and whether it happens every time. Then check the surfaces that can hold different halves of the same failure.
| Where | How to inspect it | What it can reveal |
|---|---|---|
| Browser console | Developer tools, then Console | Errors thrown by code running on the page |
| Network panel | Developer tools, then Network, then the failed request | Status, request details and the server response |
| Development terminal | The terminal running the app | Server errors that never reach the browser |
| Production build | Run the project’s real build command locally | Type, import and bundling failures hidden by development mode |
| Hosting logs | Your hosting provider’s logs | Missing configuration and production runtime failures |
| Database logs | Your database dashboard or local logs | Query, permission and row-level security failures |
Say which surfaces you checked. “The browser console fails, the request returns 500, and the development terminal shows the database error” is a chain of evidence.
If there is no error and the behavior is simply wrong, use this instead:
No error is being thrown.
I expected: [expected behavior]
I observed: [actual behavior]
Reproduction steps: [step 1, step 2, step 3]
Do not change code. Give me the three most likely places the behavior could diverge.
For each one, tell me what to inspect or log to distinguish it from the others. Do not
log secrets or personal data.
A crash stops and points at a boundary. Wrong output can keep moving through the system, which makes observation even more important.
How to grade each answer
Understand, demand a translation
You cannot judge a fix for a problem you do not understand. A useful answer names the failed operation, the value or state involved and the layer where it failed. A vague answer says there “might be an issue” and starts editing.
If the agent changes code anyway, stop it:
You changed code before establishing the cause. Restore only the change you just
made, without touching my existing work. Then answer the original question: what does
the error mean, and what evidence supports that explanation?
Narrow, compare with the last known working state
Recent changes are strong suspects, but they are not the only suspects. Environment variables, data, dependency resolution, clocks and external services can change even when the source file did not.
For a project using Git, ask for evidence such as:
Inspect the recent commit history and the diff from the last known working commit.
Also check relevant lockfile, configuration and runtime changes. Show the commands and
output you used, then give me the short list of changes that can affect this failure.
Do not change code.
A good answer is a short list tied to the failure. A dump of every file in the project has not narrowed anything.
Rank, attach evidence to every theory
One cause can be a guess wearing a confident voice. Ranking alternatives forces a comparison. The most valuable part is the check that can eliminate a theory without changing production code.
When an answer is vague, push once:
Cause number one has no evidence attached to it. Point to the exact error, log, code
path or observed value that supports it. If nothing supports it, remove it from the
ranking and tell me what evidence would be needed.
Fix one cause, keep the change coherent
Change one cause at a time, not necessarily one file at a time. Engineers call this controlling variables. If you mix two theories into one patch, a passing result cannot tell you which theory was right.
Read the proposed diff. If you do not understand it, ask:
Explain this diff section by section in plain language. For each changed file, tell me
why it is required for the proven cause and what could break if the change is wrong.
If the attempt fails, remove only that attempted change before testing the next cause. Do not discard unrelated work.
Prove, separate claims from evidence
“It works” is a claim. A reproduced flow, passing checks and inspected output are evidence. You need both halves: the original bug is gone, and the surrounding contract still holds.
After the fix is green, buy protection for the future:
Write the smallest regression test that would have failed before this fix and passes
now. Explain why it protects the reported behavior, then run it and show me the result.
When the loop does not close
Three laps with no new evidence means the current investigation has stalled. Switch strategy instead of stacking more guesses.
Strike one, audit the assumptions
List the assumptions you are making about this code path. Mark each one VERIFIED when
you have inspected evidence that proves it, or ASSUMED when you have not. Inspect the
code or runtime evidence behind the three highest-risk ASSUMED items and report what
you find. Do not change code.
Wrong fixes often grow from one confidently wrong assumption.
Strike two, compare with working
Protect any uncommitted work first. Then inspect or reproduce the last known working
commit in a separate branch or worktree. Reapply or compare changes one coherent unit
at a time until the failure appears. Engineers call this change isolation. Tools such
as git bisect can automate the comparison when there are many commits.
Strike three, shrink the problem
Create the smallest safe reproduction that still shows this bug. Remove unrelated
styling, data and features while preserving the failing behavior. Do not replace the
real implementation. If the bug disappears, tell me which removed dependency or
interaction most likely mattered and how to test that conclusion.
This is a minimal reproduction. It reduces the number of possible causes until the important difference becomes visible.
Six cases that need a different move
| Situation | Likely direction | Useful next instruction |
|---|---|---|
| Works locally, fails in production | Configuration, version, data or build difference | Compare environment variable names, runtime versions, build output and logs without printing values |
| No error, only wrong output | Logic or data divergence | Use the expected-versus-actual prompt above |
| Fails intermittently | Timing, concurrency, cache or external dependency | List every timing and ordering dependency, then add safe observations that can catch the failure |
| The agent insists it is fixed | Verification was skipped or weak | Ask for the exact output and user flow that prove the claim |
| The bug returned | The contract lacks a regression test | Add the smallest test that reproduces the original failure |
| Failure appears in untouched code | An upstream input changed | Inspect callers, shared values, resolved dependencies and runtime configuration |
Six phrases that make an agent guess
| Do not stop here | Say this instead |
|---|---|
| “Fix it” | “Here is the full error. Explain it first and do not change code.” |
| “It does not work” | “I expected X, observed Y and reproduced it with these steps.” |
| “Still broken” | “Here is the new error and output after the attempted change.” |
| “Make it work” | “Test the highest-ranked cause, then propose the smallest coherent fix.” |
| “Something is wrong” | “Here is what is wrong, where I observed it and when it started.” |
| “Just try something” | “Rank the supported causes and give me a read-only check for each.” |
A worked example
An expense dashboard loads, then turns white. The console reports
TypeError: Cannot read properties of undefined (reading 'map') at
Dashboard.tsx:42.
The first prompt translates the failure: line 42 calls .map() on a value that does
not exist at that moment. The second prompt finds three recent changes in
Dashboard.tsx, api/expenses.ts and db/schema.ts.
The ranked causes are now testable:
- The API returns
{ expenses: [...] }, but the component treats the whole response as the list. Inspect the response immediately before line 42. - The first render happens before the request finishes. Inspect the loading state and the value during the first render.
- A new user receives an empty or missing collection. Inspect the response for a new account.
The evidence confirms cause one. The coherent fix reads response.expenses, the
original flow works, nearby tests pass, and one regression test covers both the
response shape and an empty list.
That is the difference between asking for motion and running an investigation.
Make the loop a project rule
Put the durable version in the instruction file your coding tool reads. Keep the tool specific filename, but preserve the behavior:
DEBUGGING RULES FOR THIS PROJECT
- Ask for the complete error, reproduction steps and expected behavior before editing.
- Explain the failure and name its layer before proposing a fix.
- Compare with the last known working state and relevant runtime changes.
- Rank supported causes and give a read-only check for each.
- Address one proven cause with the smallest coherent change.
- Keep unrelated work untouched and remove a failed attempt before the next theory.
- Run the relevant checks and report their output before saying the bug is fixed.
- Add one regression test for a confirmed bug when the project can support it.
The file is not magic. It makes the investigation repeatable, reviewable and available in the next session.
Common questions
- Should the agent always change only one file?
- No. Start with one proven cause and ask for the smallest coherent fix. A real fix may cross a schema, API and interface, but every changed file should be necessary for that one cause.
- What if there is no error message?
- Record what you expected, what actually happened and the exact reproduction steps. Then ask where the behavior could silently diverge and what observation would distinguish those possibilities.
- How many times should I repeat the loop?
- Stop after three laps with no new evidence. Audit assumptions, compare with a known working version or shrink the problem into a minimal reproduction before trying more fixes.
- Does this loop replace debugging skills?
- No. It gives the investigation a disciplined shape. You still judge the evidence, read the diff and decide whether the test proves the right behavior.
Keep reading
20 things AI leaves out of the login it built you
The gaps AI leaves in auth, from email verification to password reset and sessions, each with a copy-ready prompt and a browser check that proves it works.
Your AI agent is stuck in a loop: 12 prompts that break it
Twelve copy-ready prompts for when your AI agent keeps saying it fixed the bug and ships the same fix, cheapest first, each with a way to tell it worked.
12 security holes AI leaves in your app, and how to close them
A copy-ready prompt to close each of the twelve security holes AI leaves in most apps, and a browser check under each one to prove it actually shut.