Skip to content

AI cost control: 13 checks before you launch

Thirteen checks that stop a runaway AI bill, each with a copy-ready prompt and a way to prove it is actually working, plus a triage plan for after.

19 min read

Most surprise AI bills are not break-ins. They are taps left running: a key shipped to the browser, a retry loop with no exit, a request with no ceiling on how much it can generate. The spending looks like ordinary usage right up until the statement arrives.

Thirteen checks close the common leaks. Each one has a prompt you can paste into your coding agent and, more importantly, a way to prove it is actually working. “I asked the agent to add rate limiting” and “rate limiting works” are different facts. A checklist item you cannot verify is a wish.

The whole list is about an hour of work. It is the cheapest hour in the project.

The thirteen checks

Work through these before anyone else can reach your app. If you only do three, do 1, 12 and 13.

The pipe, stop the leaks

  • ☐ 1. Every AI call happens on the server. No key in frontend code, ever.
  • ☐ 2. An explicit output limit on every single request.
  • ☐ 3. Rate limit per user. Ten requests a minute is a fine starting point.
  • ☐ 4. Spend cap per user, per day.
  • ☐ 5. Spend cap for the whole app, per day.
  • ☐ 6. Every loop has a hard exit. Five attempts, then stop.
  • ☐ 7. Retries use exponential backoff, not instant retry.

The meter, spend less and see it

  • ☐ 8. Identical prompts are cached.
  • ☐ 9. A small model for simple tasks. The largest model only where it matters.
  • ☐ 10. Every call logged with its token counts and the user who caused it.

The wall, what saves you when the first ten do not

  • ☐ 11. A billing alert at roughly twice your normal spend.
  • ☐ 12. A hard billing limit set in the provider dashboard.
  • ☐ 13. A kill switch: one setting that turns the AI feature off, and it is tested.

The first ten stop the leaks you predicted. The last three stop the one you did not. People skip the last three because nothing is on fire yet, which is exactly when they are cheap to build.

If the bill already arrived

Ten minutes, in this order. Do not start by reading code.

Stop the spending

Turn the feature off. If you have a kill switch, use it. If you do not, remove the API key from your server environment and redeploy, or disable the key in the provider dashboard. An app returning errors is cheaper than an app that works and bills.

Deal with the key

If you have evidence the key leaked, disable it now. The feature is already off, so there is nothing left to keep alive, and every minute the key works is a minute somebody else can spend on it.

If the key is only a suspect, rotate it in the safe order: create the new key, put it in your server environment, deploy, confirm the app works, then revoke the old one. Revoking first means downtime and a panicked deploy, and a panicked deploy is when people paste a key somewhere public and start the whole thing again.

Set the hard limit now, not later

Go to the provider dashboard and set a hard spend limit at a number you could afford to lose. Do this before you investigate. Investigation takes hours, and the tap is still open during them.

Find out where it went

Prompt
Show me the API usage logs for the last 7 days, grouped by day and by user or IP
address. Which user, key or endpoint accounts for most of the spend? Show me the top
five with their share of the total. Do not change any code.

If you have no logs, that is your first task after this. Check 10 exists because of this exact moment.

Decide what it was

Three shapes, and they need different responses.

What you find What it means What to fix
Traffic you cannot account for The key is being used by someone else Rotate it and move every call server-side
One endpoint calling itself in circles Your own code ran away Add the hard loop exit and the backoff
Many real users, each spending a little Your pricing is wrong, not your code Price the feature, then add the caps anyway

Ask, politely

Some providers will discuss a first-time runaway bill, particularly when you contact them quickly with a clear timeline. Write plainly: what happened, when you noticed, and what you have changed so it cannot repeat. Nobody owes you a refund. Ask anyway, and assume the answer is no while you fix the cause.

Then work the thirteen checks.

Checks 1 to 5, the pipe

1. Every AI call on the server

A key in frontend code is a public key. Anyone can open browser developer tools, copy it, and spend your money for as long as it works. This check is what makes the other twelve worth doing, because none of them apply to someone calling the provider directly with your key.

Prompt
Find every place this project calls an AI API. For each one, tell me whether it runs
in the browser or on the server. Move every browser-side call to a server route that
reads the key from a server-side environment variable. Then search the project and its
git history for any key, token or secret that has ever been in frontend code or
committed. Show me the list before you change anything.

Verify: open your deployed site, view the page source, and search for your key. Then open the network panel, use the AI feature, and look at the request. It should go to your own domain, never to the provider. Any key that was ever public is leaked, whatever the bill looks like, so rotate it.

2. An output limit on every request

Without a ceiling, one prompt can ask for something enormous and get it. Generated output commonly costs several times more per token than the input you send, so a single unbounded answer is often the most expensive thing in the app.

Prompt
Add an explicit maximum output length to every AI call in this project. Choose a value
per call site based on what that feature actually needs, and tell me each value with
your reasoning. If any call site genuinely needs long output, say so and cap it anyway.

Verify: ask your app for something absurd, such as a complete book, through your own interface. The answer must stop. If it stops mid-sentence, that is the check working, not a bug.

3. Rate limit per user

One script can send ten thousand requests overnight. A rate limit turns a flood into a queue.

Prompt
Add per-user rate limiting to every AI route, starting at 10 requests per minute.
Identify users by account ID when they are signed in and by IP address when they are
not. When someone hits the limit, return a clear message rather than crashing. Tell me
where the counter is stored and what happens to it if that store restarts.

Verify: use the feature eleven times quickly. The eleventh should be refused with a readable message. If nothing happens, the limit is not wired to the route you are testing.

4. Spend cap per user, per day

A rate limit caps speed, not total. Ten requests a minute is over fourteen thousand a day. The daily cap is what makes one enthusiastic user survivable.

Prompt
Add a per-user daily spend cap. Before each AI call, estimate its cost from the model
and the token counts, add it to that user's running total for the day, and refuse the
call with a clear message if it would go over the cap. Store the totals so they
survive a restart, and reset them at midnight UTC. Tell me the cap value you chose and
how a user can see how much they have left.

Verify: temporarily set the cap to something tiny, use the feature until it refuses, then set it back. A cap you have never seen trigger is a cap you do not know works.

5. Spend cap for the whole app, per day

Per-user caps assume an attacker uses one account. The global cap is what holds when a thousand fresh accounts each spend a little.

Prompt
Add a global daily spend cap for the entire app, separate from the per-user cap. When
it is reached, disable AI features for the rest of the day, show users a clear
message, and alert me immediately. Tell me the value you chose and exactly how I raise
it in an emergency.

Verify: confirm you can raise it without deploying code. A global cap you can only change by shipping is a cap that gets removed in a panic at midnight.

Checks 6 to 10, loops, waste and seeing

6. A hard exit on every loop

An agent that decides to try again has no natural stopping point. Six hours of trying costs six hours of tokens, and it happens while you sleep.

Prompt
Find every loop, retry or agent flow in this project where the model decides whether
to continue. List each one and its current exit condition first. Then add a hard
maximum of 5 iterations to each. When one hits the limit, stop, log why, and surface
the partial result rather than failing silently.

Verify: read the list it found before you accept the change. If the answer is zero loops and your app runs an agent, it did not look properly.

7. Exponential backoff on retries

When an API hiccups, a naive retry fires again immediately, and again, hundreds of times a minute. You get billed for the failures, and you make the outage worse for everyone on the provider.

Prompt
Add exponential backoff with jitter to every API retry: wait 1s, then 2s, then 4s,
then 8s, with a maximum of 5 attempts, then give up with a clear error. Only retry
errors that are actually retryable, such as timeouts, rate limits and 5xx responses.
Never retry a 400 or a 401, because those will fail identically forever.

Verify: the last clause is the one people miss. Ask which status codes are retried, and confirm that authentication and validation errors are not on the list.

8. Cache identical prompts

The same question answered twice costs twice. Many apps ask the identical thing constantly, particularly anything that summarizes or classifies fixed content.

Prompt
Add caching for AI responses. Key it on the exact prompt, the model, and any parameter
that changes the output. Tell me which call sites in this project are safe to cache
and which are not, with your reasoning. Then expose the cache hit rate somewhere I can
see it.

Verify: watch the hit rate for a day. Zero means either every prompt genuinely differs, which is fine, or your cache key contains a timestamp, which is a bug.

9. A small model for simple tasks

“Is this a valid email address”, “is this positive or negative” and “pick one of these five categories” do not need your largest model. Small models are typically far cheaper per token, and many apps spend most of their money on questions a small model answers identically.

Prompt
List every AI call in this project with what it actually asks for. For each, recommend
the smallest model that would do the job just as well, and mark the ones that
genuinely need the largest model. Then implement the routing and make the model choice
a configuration value I can change without redeploying.

Verify: run twenty real examples through both models and compare the outputs yourself. If you cannot tell them apart, the small model wins. Route by task, never by mood.

10. Log every call with its token counts

You cannot reduce a cost you cannot see, and after a surprise bill the first question is who. This check is what makes the triage section above possible.

Prompt
Log every AI call with: timestamp, user ID, endpoint, model, input tokens, output
tokens, estimated cost, duration, and whether it was a cache hit. Store it somewhere I
can query. Then build a simple internal page showing today's spend, this month's
spend, the top 10 users by spend, and the most expensive endpoint. Do not log prompt
contents or any personal data beyond the user ID.

Verify: answer this from your own dashboard in under a minute. Which user spent the most yesterday? If you cannot, the logging is not finished.

Checks 11 to 13, the wall

Most people treat these three as one thing. They are three different mechanisms, and only two of them stop anything at all.

What it does What it does not do
Alert Tells you when spending crosses a line Stop anything. It is a smoke detector, not a sprinkler
Hard limit Refuses calls past a number Help if you set it far above what you can afford
Kill switch Turns the feature off on your command Work at all if you have never tested it

11. A billing alert at twice normal spend

Prompt
What is this app's average daily AI spend over the last 14 days? Based on that, tell
me what number to set a billing alert at, at roughly twice normal, and what the alert
is called in my provider's dashboard.

Verify: send yourself a test alert if the provider supports one, and confirm it reaches somewhere you actually read at three in the morning. An alert landing in an inbox you check weekly is not an alert.

12. A hard billing limit at the provider

Every other check on this list lives in your code, which means a bug in your code can defeat it. This one lives at the provider, and it holds even when your app is completely wrong. It takes about five minutes and it is the highest-value item here.

Dashboards get rearranged, so search your provider’s own documentation for these terms rather than following someone else’s click path.

Provider Search their docs for Notes
Anthropic spend limit, usage limits Set at the organization level, and per workspace if you use them
OpenAI usage limits, budget Distinguish the notification threshold from the hard stop
Google quotas, budget alerts Cloud budgets alert by default; a hard stop is configured separately
Gateways and routers credit limit, per-key limit Usually per key, which is convenient: one key per app

There is a stronger cap that almost nobody uses. Where your provider sells prepaid credit, load a fixed amount and turn auto-recharge off. You cannot be billed for money you have not given them, and it does not depend on a limit rule firing in time or on your code being correct. The cost is that the feature dies when the credit runs out, which is exactly what you want at three in the morning and annoying at three in the afternoon. An alert at twenty-five percent remaining makes it annoying much less often.

Verify: sign in to the billing dashboard and read the number back. If you cannot find it in two minutes, it is not set.

13. The kill switch

The first twelve checks close the holes you predicted. The one that gets you is the hole nobody predicted, and it arrives at three in the morning. The switch is how you stop the bleeding in ten seconds instead of forty minutes, from a phone, without a deploy.

Prompt
Add a kill switch for all AI features, read from a single environment variable named
AI_ENABLED. Check it at the top of every AI route, before any call is made, in one
shared place that no route can bypass. When it is off, return a friendly message to
users rather than crashing, and log that the call was blocked. Then show me exactly
how to change it from my hosting dashboard without a code deploy, and how long the
change takes to take effect.

Two things make this real, and both get skipped.

Test it. Turn it off in production now, while nothing is wrong. Confirm the feature actually stops and the rest of the app does not crash. Turn it back on. An untested kill switch is a comment, not a switch.

Know how fast it is. Some hosts apply an environment variable change immediately, others need a redeploy. If yours needs a redeploy, that is not a ten-second switch, and you should back it with something faster, such as a flag row in your database that every route reads before it calls.

Verify: time yourself, from a locked phone to the feature being off. Under a minute is the goal.

Why the bill grows faster than the usage

Three pieces of arithmetic explain most surprise bills. None of them are obvious, and knowing them changes how you build.

Output costs more than input. With most providers, generating tokens costs several times more per token than reading them. A short prompt with a long answer is the expensive shape, which makes “be concise” a cost control rather than a style preference. Check your provider’s current pricing page for the exact ratio.

Conversation history is re-sent every turn. The API has no memory of its own. Turn ten sends turns one through nine again as input, so the cumulative input cost of a conversation grows roughly with the square of its length rather than in a straight line. A forty-turn chat is not four times a ten-turn chat; it is closer to sixteen. This is why long chat features surprise people. The fixes are trimming old turns, summarizing the middle of the conversation, and prompt caching where the provider offers it.

Agents multiply everything. One user action can be many model calls: plan, call a tool, read the result, decide, call again. Ten internal steps is ten times the cost of the one call you think you are making, and every step carries the growing history above. Check 6 exists because this multiplication has no natural ceiling.

Work out your own numbers rather than trusting a rule of thumb:

Prompt
For each AI feature in this app, estimate the average input tokens, the average output
tokens, the calls per active user per day, and the cost per user per month at current
prices. Show your working in a table and state which prices you used. Then tell me
which feature would cost the most if usage grew 100x, and what I should change first.

Run that before you launch and again after your first real week. The gap between the two estimates is the most useful number in the project.

Make it a project rule

Put the durable version in the instruction file your coding tool reads, so you stop having to remember any of it. Keep your tool’s filename and preserve the behavior:

Prompt
AI COST RULES FOR THIS PROJECT

- All AI calls run server-side. Never put a key in frontend code or in git.
- Every AI call sets an explicit maximum output length.
- Every AI route is rate limited per user and checks the per-user and global daily
  spend caps before calling.
- Every loop where the model decides to continue has a hard maximum of 5 iterations.
- Retries use exponential backoff with jitter, at most 5 attempts, and only on
  timeouts, rate limits and 5xx responses.
- Cache identical prompts. Never include a timestamp in a cache key.
- Route each task to the smallest model that does it. Model choice is configuration,
  not code.
- Log every call: user, model, input tokens, output tokens, estimated cost, cache hit.
  Never log prompt contents or personal data.
- Every AI route checks AI_ENABLED before doing anything, in one shared place.
- When you add a new AI call, say in your summary which of these rules you applied.

The file is not magic. It makes the rules repeatable, reviewable, and available in the next session instead of only in your head.

Audit an app you have already shipped

Prompt
Audit this project for AI cost risk. For each of these 13 checks, tell me whether it
is fully in place, partly in place, or missing, and point at the file and line that
proves your answer: server-side calls only, an output limit everywhere, per-user rate
limit, per-user daily spend cap, global daily spend cap, hard loop exits, exponential
backoff, prompt caching, model routing by task, per-call logging with token counts,
billing alert, hard billing limit, kill switch. Do not fix anything yet. Rank what is
missing by how much money it could cost me this week.

Take the top of that ranked list and work down. Most apps are three fixes away from safe, and the three are usually the same three: the call is in the browser, nothing caps the output, and the provider has no limit set.

Common questions

Is a billing alert enough on its own?
No. An alert tells you that spending crossed a line; it does not stop the spending. Pair it with a hard limit at the provider and a tested kill switch, so something refuses calls even while you are asleep or your own code is wrong.
Which check should I do first if I only have five minutes?
Set the hard spend limit in your provider dashboard. It lives outside your code, so it still holds when a bug defeats every safeguard you wrote yourself. Moving API calls to the server and testing a kill switch come next.
My key was in frontend code but the bill looks normal. Do I still need to rotate it?
Yes. Anything shipped to a browser is public, and normal spending only means nobody has used it yet or nobody has used it noticeably. Rotate the key, move the call server-side, and check your usage logs for traffic you cannot account for.
Will my provider refund a runaway bill?
Nobody owes you a refund, and you should not plan around getting one. Some providers will discuss a first incident if you contact them quickly with a clear timeline and the changes you made. Ask politely, then assume the answer is no.