I write most of my code with AI these days. Claude Code, Codex, whatever fits the job. I'm not going back. It's faster and most of the time the code is fine.
But "fine" means it does what I asked. The model builds the feature you describe. It doesn't sit there wondering what someone with bad intentions will do with it, unless you ask. And most people don't ask.
So here's what I'd check first in any app that was built mostly by AI. None of this is new. These are old mistakes, but AI makes them faster and at a lot more volume.
1. Secrets in the frontend
This is the big one. You tell the model to "connect to Stripe" or "call OpenAI from the dashboard," and the fastest working version puts the key right in the frontend code. Frameworks make it easy to do by accident. Anything prefixed NEXT_PUBLIC_ or VITE_ gets bundled into JavaScript that every visitor downloads.
How to check: build your app, then search the output for things like sk_live, sk-, service_role, or long strings starting with eyJ. If you find a secret key in there, assume it's already been copied. Rotate it, then move the call to the server.
2. Auth that only lives in the UI
The model hides the "Delete user" button for non-admins. Great. But the API endpoint behind that button doesn't check anything, so anyone who finds it in the network tab can call it directly.
A hidden button is not access control. Every API route needs to check who's calling and whether they're allowed to do what they're asking. Test it by logging in as a normal user, copying a request from dev tools, and pointing it at something that user shouldn't be able to touch.
3. Changing the ID in the URL
/api/invoices/1042 returns your invoice. What does /api/invoices/1043 return? If it's someone else's invoice, you have what's called an IDOR (insecure direct object reference). It's one of the most common serious bugs out there, and generated code is full of it, because the query to "get invoice by ID" works perfectly in every demo.
The fix is boring: every lookup should also filter by the current user or their org.
4. Database rules left wide open
If you're on Supabase or Firebase, your frontend talks to the database directly and security rules are the only thing stopping people from reading every table. The model will often get things working by turning Row Level Security off, or it'll leave Firebase in test mode because that's what the quickstart did.
Go into the dashboard and look at every table. If RLS is off on anything with user data, that data is public.
5. Error messages that say too much
Stack traces, SQL errors, internal file paths, framework versions. Generated code tends to pass errors straight back to the client because it makes debugging easier. It also hands an attacker a map of your stack. In production, users get a generic message and the details go to your logs.
6. No rate limits, especially on AI endpoints
Login, password reset, one-time codes, signup. Without rate limits, someone can guess passwords or codes all day.
And if your app has an AI feature, this gets expensive fast. An endpoint that calls a model with no auth and no rate limit is basically a free API for anyone who finds it, and you're paying the bill. I work on LLM infrastructure in my day job and I'd put this near the top of the list. Cap it per user, and put a hard spending limit on your provider account.
7. Packages nobody checked
Models sometimes suggest packages that are old, abandoned, or don't exist at all. That last one is a real attack now. People register the package names models tend to make up and put malware in them. Before you install something the model suggested, check that it's real, maintained, and actually the one you meant.
What a pen test does and doesn't catch here
A black box pen test looks at your app from the outside, the way an attacker would, so it's good at finding the stuff that's visible from there: exposed secrets, verbose errors, missing headers, weak cookie settings, known vulnerable components. Depending on what's reachable, it can also turn up broken access control.
It won't read your code, so it's not a replacement for looking at your own database rules and API checks. The best setup is both. Go through this list yourself, then get someone independent to poke at it from the outside.
If you want a quick first look, the free header scan takes about ten seconds. If you're about to put an AI-built app in front of real customers, a full pen test is $495 and back in 24 hours.