Defending your codebase from AI slop: guardrails that keep you in control
/ 10 min read
AI coding agents are fast. They also write code that compiles, looks fine at a glance, and slowly wears down your architecture. A component imports the database layer directly. A function grows to 200 lines with three nested ternaries. A dependency shows up that nobody asked for. A routes file picks up a formatDate helper, then a slugify, then a retryWithBackoff. Tests run every line but never check a result.
None of this is new. Junior developers in a hurry do the same things. What’s new is the volume. If an agent opens ten PRs a day, careful human review alone won’t catch everything.
The answer isn’t to stop using AI. It’s to write your standards down in a form machines can enforce, so the rules hold no matter who or what wrote the code. Here’s the stack I usually build for my projects.
1. Write the rules down first
Tools can only enforce rules that exist. Before adding any tooling, I decide how I want my code to look, define the rules, and write them in two files:
ARCHITECTURE.mdcovers the layers, which layer may import which, where state lives, and how data flows.AGENTS.md, essential nowadays, is usually created automatically and then edited by me. It covers conventions for AI agents: naming, patterns to use, patterns to avoid, and tools that are off-limits (for example: “we use Biome, don’t add ESLint”).
Agents read these files. So do new hires. And every tool below points back to a section in one of them, so when a check fails, the message explains why and not just what.
2. Guard the boundaries: dependency-cruiser
Something I’ve discovered quite recently: dependency-cruiser checks which files import which. You describe the allowed directions between layers, and it fails the build when something crosses a line:
- UI components can’t import the database layer
- modules can’t import each other circularly
- features can’t reach into another feature’s internals
- orphaned files that nothing imports get flagged
This is the kind of mistake AI makes most, because it takes whatever import path is closest. A rule that always gives the same answer stops it before review.
3. Guard the shape of the code: Semgrep
dependency-cruiser tells you which files talk to each other. It can’t tell you what the code inside a file is doing. That’s what Semgrep is for.
How Semgrep works
Semgrep is a pattern matcher that understands code. You write a rule that says “code shaped like X, in files matching Y, is a violation.” Semgrep parses each file into a syntax tree and reports every match. It doesn’t run your code or try to understand your whole app. It only checks shapes.
A rule is a short YAML entry:
rules: - id: api-routes-no-helpers languages: [python] severity: ERROR message: > routes.py is for route handlers only. Move helpers to their own module (see ARCHITECTURE.md → API). paths: include: [packages/api/src/api/routes.py] patterns: - pattern: | def $FUNC(...): ... - pattern-not-inside: | @$ROUTER.$METHOD(...) def $F(...): ...patternis the code to match.$FUNCmatches any name, and...matches anything.pattern-not-insideis the exception. Here, functions with a route decorator are allowed.pathslimits the rule to certain files.messageis what the developer (or the AI agent) sees when the rule fails. Write it as an instruction.
Running semgrep --config .semgrep/rules/ checks the whole repo against every rule in that folder. We keep one YAML file per area (app, API, mq) so rules stay easy to find, and each rule gets a small test fixture with code that should match and code that shouldn’t, so we know the rule catches what we meant.
Read their docs (or ask your agent to) on what’s possible and get creative.
One-purpose files
This was the rule we wanted most. Some files exist for one job, and AI agents love to drop “just one small helper” into them. After a few months, your routes file is half utilities.
So we made it impossible. Each of these files may only contain its one kind of thing:
- API route modules (
routes.py): route handlers only - App
services/: one class, methods only, no loose functions - App
functions/: TanStack server functions only - Hooks: hooks only
- Route files: route definitions only
- DB schema: table definitions only
- API models/schemas: models only
- API repositories: repository classes only
Helpers go in their own module, where they can be named, tested and reused. The rule message tells the agent exactly that, so it usually fixes itself on the next try.
Other patterns worth enforcing
Once you think in shapes, a lot of review comments turn into rules:
- query and mutation keys written inline instead of coming from a key factory
- data fetching inside components
- raw
fetchin routes and server functions - Drizzle queries outside
db/queries - exported query functions that default to
tx: DbClient = db(this hides transaction bugs) - queue names written as string literals
- raw SQL built from strings in the Python API
- API routes importing models directly instead of going through a repository
Semgrep works on many languages, so it covers anything you want.
Rule of thumb: if a review comment can be written as a code pattern, make it a Semgrep rule. Then it gets checked the same way on every PR.
Adding it to an existing codebase
You probably have violations already. Don’t let that stop you. In CI, Semgrep can compare against a baseline:
- On PRs, only violations added in that PR fail the check.
- On
main, the full scan runs, so you can see the existing debt and pay it down over time.
New code is held to the standard from day one, and nobody has to fix everything first.
4. Scan for vulnerabilities: SAST, SCA and DAST
Architecture rules keep the code tidy. They don’t tell you whether it’s safe. AI agents write the same security bugs humans do, just faster: SQL built from strings, an endpoint that forgets to check who’s calling it, a token pasted into a config file, user input rendered as HTML.
Security tools fall into three groups, and each sees something the others can’t:
| Kind | Looks at | Finds | Misses |
|---|---|---|---|
| SAST (static application security) | Your source code, not running | Injection, unsafe APIs, hardcoded secrets, tainted data flows | Runtime config, auth bugs that depend on real data |
| SCA (software composition analysis) | Your dependencies and lockfiles | Known CVEs, malicious packages, licenses | Bugs in your own code |
| DAST (dynamic application security) | The running app, from outside | Missing auth, exposed endpoints, bad headers, misconfigured servers | Where in the code the bug is; code paths it never reaches |
SAST and SCA run on every PR in seconds or minutes. DAST needs a deployed app, so it runs against a preview environment or staging. Most of the tools below cover more than one group:
| Tool | SAST | SCA | DAST | Other |
|---|---|---|---|---|
| Semgrep | ✅ | ✅ | Secrets | |
| Trivy | ✅ | Containers, IaC, secrets | ||
| Snyk | ✅ | ✅ | ✅ | Containers, IaC |
| Checkmarx One | ✅ | ✅ | ✅ | Containers, IaC, secrets, API |
| GitHub CodeQL | ✅ | Dependabot covers SCA | ||
| ZAP | ✅ | Free, open source |
I am an OSS supporter, so I usually go with Semgrep and Trivy. Companies I’ve worked for went with whatever suited them best. Choose any; no judgement there.
SAST: your code
Good SAST tools track data flow: they follow user input from a request handler and flag it if it reaches SQL, a shell or HTML unescaped. Ask an agent to “add search” and you may get:
@router.get("/search")def search(q: str, db: Session = Depends(get_db)): return db.execute(text(f"SELECT * FROM products WHERE name LIKE '%{q}%'"))It works, and it’s SQL injection. Semgrep Code catches it with the engine you already run for architecture (add --config p/owasp-top-ten). CodeQL is free for public repos and shows results in the PR. Snyk Code and Checkmarx cover the same ground in their platforms. Fail the build on high-confidence rules only. A noisy check gets switched off.
SCA: your dependencies
Agents add packages freely, sometimes outdated, sometimes made up and squatted (“slopsquatting”). Trivy is the free baseline for vulnerable packages, secrets, container images and IaC:
trivy fs . --scanners vuln,secret,misconfig --severity HIGH,CRITICAL --exit-code 1Trivy only matches versions, so it flags vulnerable packages whether or not you use the vulnerable part. Semgrep Supply Chain checks reachability: does your code actually call the vulnerable function? Snyk opens fix PRs and alerts when a new CVE hits code you already shipped. Checkmarx One also detects malicious packages such as typosquats.
Require CODEOWNERS approval on package.json and lockfiles, so every new dependency is a human decision.
DAST: the running app
DAST sends real requests to a deployed app, so it catches what code scanning can’t: an endpoint the agent left without an auth check, one user reading another’s data by changing an ID, missing security headers, exposed debug pages. Run it against a preview deploy or staging:
docker run -t ghcr.io/zaproxy/zaproxy:stable zap-api-scan.py \ -t https://preview.example.com/openapi.json -f openapiZAP is free. Nuclei checks for known exposures. StackHawk, Snyk API & Web and Checkmarx DAST are paid options. Give the scanner credentials and an API spec, or it will scan your login page and little else.
DAST, I’d say, is a nice-to-have, not a must-have. Folks at Semgrep think so too. Occasional manual testing or pentesting is more than enough.
5. Measure complexity and test quality
Something I’ve learned about from Uncle Bob. I never thought about it until AI became a thing.
CRAP stands for Change Risk Anti-Patterns, a metric used to identify risky, complex, and poorly tested code. The CRAP Index measures the maintenance risk of a specific function or method.
Rule of thumb: a score above 30 indicates a high-risk, “CRAPpy” method that needs attention or refactoring.
I found two tools I could integrate into my pipelines that would help to keep AI-generated code under control:
- crap4ts computes a CRAP score from complexity and coverage. Complex code with weak tests scores high. Set a limit and fail functions that go over it.
- ArchUnitTS lets you write architecture rules as unit tests, if you’d rather keep them next to your test suite.
6. AI reviewing AI: CodeRabbit with custom rules
Some rules can’t be written as patterns: “does this name match what the code does?” or “is this abstraction premature?” For those, we can use CodeRabbit (or one of many other AI code review tools) with a strict config file, like .coderabbit.yaml:
- the strictest review profile, and it requests changes when it finds problems
- it reads
ARCHITECTURE.mdandAGENTS.mdas guidelines - instructions for each folder, each linked to the doc section it comes from
- pre-merge checks: most only warn, but a few block the merge (for example, a schema change with no migration)
The important part is that CodeRabbit is the last layer, not the first. AI review isn’t consistent, so anything that can be checked by a tool that always gives the same answer should be. CodeRabbit handles judgement calls.
7. More you can add
What we’ve set up covers most of it. Here are a few more tools that can help you keep even tighter control over your repo:
- Knip finds unused files, exports and dependencies. AI leaves a lot of dead code behind.
- Mutation testing (which Uncle Bob is also looking at and seems to think is an essential part of CI/CD checks) (Stryker) changes your code on purpose and checks whether tests fail. It’s the best way to catch tests that run code but never check results.
- PR size limits: fail or warn on PRs over a few hundred lines. Small PRs actually get reviewed. Not something I’d do, though.
- Pre-commit hooks: such as Lefthook or, most commonly in JS-based projects, husky. You can configure them to run anything, including Biome, Semgrep and type checks, locally, so you get feedback before you push.
The principle
Every layer follows the same idea: move each rule to the most reliable tool that can enforce it.
- Write the rule in
ARCHITECTURE.md/AGENTS.md. - If it’s about imports → dependency-cruiser.
- If it’s about code shape → Semgrep.
- If it’s a security bug in your code → SAST (Semgrep, CodeQL, Snyk Code, Checkmarx).
- If it’s about dependencies or secrets → SCA (Trivy, plus Semgrep Supply Chain, Snyk or Checkmarx).
- If it only shows up at runtime → DAST (ZAP, Nuclei, StackHawk, Snyk, Checkmarx) against a preview deploy.
- If it needs judgement → CodeRabbit, and then a human.
AI can write as much code as it likes. It just has to pass the same checks as everyone else. That’s how you stay in control.