7 MIN READ
Picture the pull request that lands on a Tuesday afternoon. The diff looks tidy. The variable names are sensible. The function does exactly what the ticket asked. The merge runs green, the deploy ships, and two weeks later the on-call gets paged because the code that looked clean in the diff turned out to be a textbook SQL injection. The code was AI generated. Nobody on the receiving end had any way to know. That gap is the gap most engineering groups are quietly learning to close in 2026.
Most engineering groups shipped their first AI generated PR sometime in 2023. By 2026 the practice has moved from “experiment” to “default workflow,” and the honest conversation has shifted from whether to allow it to how to review it. The useful question is no longer whether the developer should use AI generated code. They already do. The question is whether the review practice treats the output with the same discipline it would apply to a pull request from an unfamiliar contractor who has never seen the codebase. The orgs that have answered that question well have a short list of patterns they look for. The ones that have not are finding the patterns the hard way, in production.
What the model does and does not know
Start with the easy one. The model has read more code than any human reader will ever see. It knows the shape of a Python try/except, the idiomatic way to paginate a Postgres query, the standard library function for parsing a URL. That knowledge is the reason the diff looks tidy in the first place. Treating the model as a junior developer with a perfect memory of Stack Overflow is closer to the truth than treating it as a senior engineer who understands the system.
Here is the harder part. The model does not know the codebase. It does not know that the authentication middleware wraps every route handler in a check the controller code never has to repeat. It does not know that the data access layer already strips the soft deleted rows before they reach the service. It does not know which functions in the legacy module are about to be deleted next sprint. The model produces code that satisfies the prompt. The prompt is rarely the whole story, and the gap between what the prompt asked for and what the surrounding system actually needs is where the production incidents come from.
Practical upshot. Treat every AI generated PR as if it came from a contractor who has read the docs but has never run the code. The review has to catch what the model did not know to ask about, which means the person doing the review has to know what to ask about in the first place. The orgs that get this right have made that part of the checklist rather than a question of individual skill.
Where the code actually breaks
Most of the failure modes land in the same handful of buckets. The pattern is the pattern every security group has been warning about for years. The new bit is that the code looks clean on the first read, which means the pattern is harder to spot than it would be from a less polished diff.
Authentication and authorisation boundaries. The model produces code that checks the user is logged in. It rarely produces code that checks the user is allowed to do this specific thing to this specific resource. The authn vs authz gap shows up in almost every AI generated PR that touches a privileged endpoint, and the gap stays invisible until someone tests the endpoint as a user who is logged in but should not have access. The fix is a step in the review that explicitly walks the authorisation model for the touched endpoint, not just the authentication.
Input validation. The model produces code that handles the happy path and the obvious edge cases. It does not always produce code that handles the inputs a real attacker would actually send. Null bytes in file paths, Unicode normalisation in search fields, integer overflow in pagination cursors, all the places a working security group already has tests for. The PR passes the unit tests. It does not pass the fuzzer, because nobody thought to run the fuzzer on a PR that was so clean looking on the first pass.
Error handling. The model produces code that throws on the failure path. It does not always produce code that handles the failure in a way the system can actually recover from. Empty catch blocks, swallowed exceptions, generic 500 responses that hide the real failure from the operator, all of it shows up in AI generated diffs at a higher rate than it shows up in human written diffs. The reviewer has to specifically look for it because the happy path looks fine.
Concurrency. The model produces code that works in the single threaded test. It does not always produce code that works when the same row gets updated twice in the same millisecond from two different workers. Race conditions, lost updates, deadlocks in the migration script, the long tail of bugs that only show up under load. These are the bugs the integration test suite would have caught if the integration test suite had been written for the new code path, which it usually has not, because the new code path was AI generated and shipped before anyone had time to think about the load profile.
The review practice that holds up
The orgs that have stopped finding AI generated bugs in production share a few habits. None of them are exotic. All of them require the AppSec lead, the engineering manager, and the second pair of eyes to agree on the list, and all of them require the list to be applied on every AI generated PR, not just the ones that look suspicious on the first read.
The first habit is a threat model attached to the PR description, not buried in the system architecture document nobody reads. Two or three sentences on what the changed code is allowed to do, who is allowed to call it, and what the worst case looks like if it gets called by someone who should not. The PR template enforces it, the AppSec lead audits the audit trail. The threat model is the document that makes the rest of the review faster, because nobody has to reconstruct it from the diff.
The second habit is a separate review pass for AI generated code, run by someone who is not the same person who wrote the code. The engineer who is going to ship the feature is the wrong person for the security pass, because they have a vested interest in the PR going through. The rotation does not have to be slow or formal. A second pair of eyes that knows the threat model and is willing to push back is the difference between a 10 minute review and a four hour postmortem.
The third habit is a test that fails when the code does the wrong thing, not just a test that passes when the code does the right thing. Property based tests, fuzz tests, integration tests that hit the new endpoint with the inputs the production traffic actually carries. The test is what catches the failure mode the human reader missed, and the test is what the next PR can run to confirm the next AI generated change did not regress the same edge case.
What this means for the org shipping it
None of this requires buying a new tool. None of it requires hiring a new role. The platform org has to enable it. The security org has to bless the threat model template. The engineering manager has to make the second review pass the default for any PR tagged as AI generated. The procurement lead has to stop paying for “AI code review” features that just run the model against itself, which is the AI equivalent of asking the contractor to mark their own homework.
Honest framing for 2026. AI generated code is now a normal input to the production system. The review practice that handles that input well looks like the review practice that has always handled unfamiliar contractor code well. The threat model in the PR. The second pair of eyes. The tests that catch the failure modes. The work has not changed. The discipline has.

The bottom line
AI generated code is in the production stack. The diff is going to look clean. The person on review duty is going to have to assume the model did not know the codebase and check for the four failure modes anyway. The org that treats the threat model, the second review pass, and the failing test as the default for every AI generated PR is the org that ships AI generated code without finding out about the SQL injection in production.
Sources & Further Reading
All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.
Spotted an error? Email the editor. Corrections are issued with a visible correction note.
Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.



