Skip to content

How to review AI code without reading it: 10 questions, 7 checks

Ten questions that make AI-written code confess, how to tell a real answer from a confident one, and seven hands-on checks. Two of them find data leaks.

15 min read

You do not need to read code to review it. Experienced engineers do not read every line of every change either. They ask questions until the code gives itself away.

The catch nobody mentions is that questions only work if you can tell a real answer from a fluent one. An AI will answer every question you ask, confidently, in good English. A question you cannot grade reassures you instead of informing you, which is worse than not asking at all.

So this guide gives you three things: ten questions to ask the AI after every feature, how to grade each answer (what a real one sounds like, what a dodge sounds like, and the exact follow-up to paste when you get the dodge), and seven checks you do with your own hands, two of which find the kind of flaw that leaks other people’s data.

Nothing here assumes you can read code. There are no snippets to understand and no syntax to learn. The whole review takes about fifteen minutes, and ten of those are the AI working.

The review checklist

Print this, or keep it in a note. The rest of the article explains how to use it.

Ask the AI these, after every feature

  1. Explain this file like I’m 12.
  2. What did you assume that I never told you?
  3. What packages did you add, and why each one?
  4. What happens if a user does the wrong thing here?
  5. Where exactly do you check that this user is allowed to see this?
  6. What would an attacker try first on this page?
  7. Show me the 3 riskiest lines you wrote.
  8. What did you NOT test?
  9. If this breaks, where will the error show up?
  10. What would a senior engineer change about this?

Then close the chat and do these yourself

  • ☐ 1. Click every button.
  • ☐ 2. Type garbage into every form: an emoji, 5,000 characters, and nothing at all.
  • ☐ 3. Serious: log out, then try to open a page you should not be able to see.
  • ☐ 4. Serious: change the ID in the address bar to another account’s record.
  • ☐ 5. Slow your connection down and watch what loads.
  • ☐ 6. Open it on your phone.
  • ☐ 7. Do the main task three times, fast.

The two marked serious are not polish. They are the ones that find data leaks. Do those two even if you skip everything else.

And the rule under all seventeen: if something feels wrong, it probably is. Paste what you saw back into the chat and ask why.

How to tell a real answer from a confident one

Learn these three tells and you can grade a question that is not even on the list.

Tell 1: specific beats general. A real answer points at something: a file, a line, a value, a name. “The authentication is handled securely” is not an answer, it is a mood. “The check is in the middleware file, line 40, and it only runs on pages under /app” is an answer. You did not need to read a line of code to hear the difference.

Tell 2: an interrogation with no admissions has failed. Ten honest questions about real software usually produce a few “I didn’t”, “I’m not sure” or “I assumed” replies. If every answer comes back confident and clean, you are probably not being told about the code. You are being told what you want to hear. Ask question 8 again, harder.

Tell 3: the confident hedge. “Should be fine”, “typically”, “generally”, “in most cases”, “best practice suggests”. Words like these often mean it did not go and look. They are the most useful words in this guide, because you can spot them without understanding anything around them.

When any answer feels smooth and empty, one follow-up works on all ten questions:

Prompt
Point at the exact file and line that makes you say that. If you did not
actually check, say so instead of guessing.

Giving the AI explicit permission to say it does not know makes it say so more often. That second sentence is doing most of the work.

Accept one thing before you start: you are not looking for a clean report. A feature that produces three honest problems has been reviewed. A feature that produces none has only been described.

Questions 1 to 5: what it does, and who can see it

1. Explain this file like I’m 12

Why it works: anything genuinely understood can be said simply. The question tests the explanation and the code at the same time, and it costs nothing.

A real answer: plain sentences about what happens, in order. “When someone submits the form, it saves the order, then it sends the confirmation email.”

A dodge: jargon you cannot follow, or a description of what the code is rather than what it does.

Follow-up:

Prompt
Say that again with no technical words at all. If you cannot, tell me
which part you do not fully understand yourself.

2. What did you assume that I never told you?

Why it works: this is the highest-yield question on the list. Many problems in AI-built software are not mistakes. They are reasonable guesses about things you never specified, and they stay invisible until a real user does something you did not picture.

A real answer: a short list of specific guesses. “I assumed prices are in one currency. I assumed one person edits a record at a time. I assumed email addresses are unique.”

A dodge: “I didn’t make any assumptions.” That is almost never true.

Follow-up:

Prompt
Every build makes assumptions. Give me five, even small ones, and mark
which ones would be expensive to change later.

3. What packages did you add, and why each one?

Why it works: every added package is code you did not write, cannot see and now depend on. Packages are not bad. The question is whether each one does enough to be worth depending on.

A real answer: each name with a one-line job, and an honest note where one is doing very little.

A dodge: a list with no reasons, or a package that is there “for utilities”.

Follow-up:

Prompt
Which of these could I remove with less than an hour of work? Is any of
them barely used? Do not remove anything yet.

4. What happens if a user does the wrong thing here?

Why it works: the feature was most likely built for the path where everything goes right. Real users leave that path all the time, and they are not being difficult when they do.

A real answer: specific wrong moves and what currently happens for each. “Empty form: shows an error. 5,000 characters: saves it all and breaks the layout. Double-click on submit: creates two orders.”

A dodge: “There is validation in place.”

Follow-up:

Prompt
List ten specific wrong things a real person could do on this screen.
For each one, tell me what the app does today, not what it should do.

5. Where exactly do you check that this user is allowed to see this?

Why it works: this is the question that finds the serious problem. In apps built quickly, AI-assisted or not, one of the most common serious flaws is not a crash. It is a page or a record that anyone can reach if they know the address. The word doing the work in this question is exactly.

A real answer: a specific place, and honesty about what it covers. “In the middleware, which covers every page under /app. The API routes check separately, and two of them do not.”

A dodge: “The user has to be logged in.” Logged in is not the same as allowed. Everyone who signs up is logged in.

Follow-up:

Prompt
Being logged in is not the same as being allowed. For each page and each
API route, tell me whether it checks that this specific user owns this
specific record. List the ones that do not. Do not fix anything yet.

That follow-up is the single most valuable prompt in this guide. Run it before any launch.

Questions 6 to 10: what breaks, and what was skipped

6. What would an attacker try first on this page?

Why it works: it moves the point of view from builder to attacker, and it tends to produce a list you can test by hand in a few minutes with no security knowledge.

A real answer: concrete, boring, ordinary attempts. Changing a number in the address bar. Submitting the form while logged out. Submitting the same form many times in a row.

A dodge: a lecture about general security principles.

Follow-up:

Prompt
Give me the five simplest ones: things I could try myself on my own app,
in a browser, in five minutes, with no tools. Number them.

Try them only on your own app, with test accounts you created.

7. Show me the 3 riskiest lines you wrote

Why it works: it forces a ranking. “Is this safe?” invites reassurance. “Which parts are weakest?” gets you a list, because you asked it to compare rather than to judge.

A real answer: three specific places, each with a reason and a note on what breaks if it is wrong.

A dodge: “Nothing here is particularly risky.”

Follow-up:

Prompt
Something is always the riskiest. Rank them anyway, even if all three are
low risk, and tell me what breaks if each one is wrong.

8. What did you NOT test?

Why it works: it asks for the negative space, which nobody volunteers. The gap between “the tests pass” and “this works” lives inside this answer.

A real answer: an honest map. “I checked the normal path. I did not check what happens with no data, with a very large amount of data, or when two people do this at the same time.”

A dodge: “The implementation follows best practices.” That does not answer the question.

Follow-up:

Prompt
List what is untested, ordered by how likely a real user is to hit it in
their first week.

9. If this breaks, where will the error show up?

Why it works: it is the only question here that helps you after you ship, and it is the one people skip. Knowing where to look is the difference between a five-minute fix and a lost evening.

A real answer: a named place. The browser console, your terminal, your hosting provider’s logs, your database’s logs. Ideally with roughly what the message will say.

A dodge: “It will be logged.”

Follow-up:

Prompt
Tell me step by step how I would find that error. Then tell me whether
anything here could fail silently, with no error anywhere at all.

The second half matters more than the first. A silent failure is worse than a crash, because a crash stops and a wrong number ships.

10. What would a senior engineer change about this?

Why it works: it invites the criticism an AI tends not to volunteer while it is presenting finished work. Asking it to speak as someone else often loosens the answer.

A real answer: two or three specific changes, each with why it matters and roughly how much work it is.

A dodge: generic advice about tests and documentation, with nothing about your actual feature.

Follow-up:

Prompt
Of those, which one would you do first if you had one hour, and which are
fine to leave forever? Be honest about which ones do not matter for an
app this size.

That last sentence matters. Not every suggestion is worth doing, and a review that produces twenty must-dos tends to get ignored entirely.

The seven checks you do with your own hands

The AI can be wrong about its own work. Your hands cannot. These take about five minutes and need no technical knowledge.

They are not equally important. Two of them find data leaks, and the rest find rough edges. If you only ever do two, do the two marked serious. Run them on your own app, with test accounts you created, and use your payment provider’s test mode for anything involving money.

# Do this You are looking for If it fails
1 Click every button and link, including the ones you never use A button that does nothing, or the wrong thing Rough edge. Fix it before a demo
2 In every form, try one emoji, then 5,000 characters, then submit it empty A crash, a broken layout or bad data saved Rough edge, unless it saves garbage
3 Log out, then paste the address of a page only logged-in users should see The page opening anyway Serious. Anyone can read it
4 Log in as one test account and change the ID in the address bar to one owned by another Someone else’s data appearing Serious. This is a data leak
5 Slow your connection down and reload Something that never finishes and never says so Rough edge, but users will think it is broken
6 Open it on your actual phone, not a narrowed browser window Things off-screen, or hidden under the thumb bar Rough edge, and often the first thing a client notices
7 Do the main task three times, quickly Duplicate records, or a double charge Serious if money or records are involved

For check 5, desktop browsers can simulate a slow connection from their developer tools. Search your browser’s help pages for “network throttling”.

Why checks 3 and 4 matter most

This category of flaw has a name: broken access control. It sits at number one in the OWASP Top 10, the most widely cited list of web application security risks.

It happens because “is this person logged in?” is one check, written once, and “is this person allowed to see this specific thing?” is a separate check that has to be written everywhere. It gets missed in one or two places. You do not need to understand the fix to find it. You need thirty seconds and the address bar.

When a check fails, get the full list first

Paste what happened back into the chat exactly, and ask for the scope of the problem before any fix:

Prompt
I logged out and opened [paste the address] and the page loaded anyway.
Do not fix it yet. Tell me which other pages and API routes have this
same problem. List all of them first.

A leak in one place is rarely in only one place. Fixing them one at a time as you stumble on them is how you end up shipping with three still open.

Make the AI review itself without being asked

Paste this block into your coding tool’s project instruction file, such as CLAUDE.md, AGENTS.md or your Cursor rules. If you do not have one yet, the 5-file starter pack explains what it is and where it goes.

Prompt
AFTER EVERY FEATURE, BEFORE I ASK
End every completed feature with a short review section:
- What you assumed that I never told you.
- What you did NOT test.
- The 3 riskiest things you wrote, ranked, with what breaks if
  each one is wrong.
- Exactly where you check that this user is allowed to see this
  data, page by page and route by route.
- Where an error from this will show up, and whether anything here
  can fail silently.
Never say "done", "it works" or "secure" without evidence. If you did
not check something, say so instead of guessing.

That turns the checklist into something that happens whether you remember it or not, which is the point. You still grade the answers, and you still do the hand checks.

When to run which part

Moment What to run
After every feature Questions 1, 2, 4, 5 and 8, plus checks 1 and 2
Before showing anyone All ten questions and all seven checks
Before launch All seventeen, plus the question 5 follow-up on every page and route
Whenever the AI says “done” or “it works” Question 9, every time

What this review does not catch

This is a review, not an audit, and knowing its edges is part of using it well. It will not reliably find problems that only appear under real load, things that break when two people act at the same moment, known vulnerabilities inside the packages you installed, or business logic that is technically correct and wrong for your business.

That last one only you can find, because only you know what the numbers are supposed to say.

When to stop reviewing and bring in a person. If you take payments, store health or identity documents, hold data that belongs to other companies, or pass a few hundred real users, have an experienced engineer look at it at least once. At that point a professional review costs far less than cleaning up a leak. Everything in this guide still applies. It just stops being the whole plan.

The seventeen moves matter less than the habit. You are not proving that the code is perfect. You are refusing to ship something you have never questioned. That refusal is the skill, and it does not require reading a single line.

Common questions

Can I review AI-generated code if I cannot read code?
Yes, if you review by questioning rather than reading. Ask the AI specific questions about what it assumed, who can see what and what it did not test, then grade each answer by whether it points at something specific. Finish with checks you do by hand in the browser, which need no technical knowledge at all.
What is the single most important check for an AI-built app?
Log out and open a page only logged-in users should see, then log in as one test account and change the ID in the address bar to a record owned by another. Both test for broken access control, which OWASP ranks first in its Top 10 list of web application security risks. Each takes about thirty seconds.
Can I just ask the AI to review its own code?
It helps, but it is not enough on its own. The AI can be wrong about its own work and can sound confident while it is. Ask it to point at the exact file and line behind every claim, give it permission to say it did not check, and confirm the serious answers with your own hands in the browser.
When do I need a human engineer to review my app?
When you take payments, store health or identity documents, hold data that belongs to other companies, or grow past a few hundred real users. At that point a professional review costs far less than cleaning up a leak. Everything in this checklist still applies; it just stops being the whole plan.