Skip to content
▌beetlix/swarm
← All reviews

AI Written Code Guide 2026: Ship Faster, Safer

4.2/ 5
Arif AriyanReviewed by Arif Ariyan · Senior Software Engineer ·

What AI written code actually looks like in 2026

The honest answer is that most of it is fine, and the part that is not fine is expensive. GitHub's Octoverse reporting has tracked AI-assisted contributions climbing year over year, and the 2024 edition put Copilot-assisted work at a meaningful share of code written on the platform. By 2026 the question is no longer whether AI written code is in your repository. It is. The question is whether your review process knows the difference between the boilerplate and the auth middleware.

I want to be precise about what "AI written code" means here, because the phrase gets used to cover three different things. There is code an assistant autocompleted one line at a time. There is code an agent generated from a prompt across multiple files. And there is code a human wrote and then asked a model to refactor. The failure modes differ across all three, and so does the review burden. A 40-line autocomplete suggestion inside a function you already understand is a different risk object than a 600-line agent-generated module that touches your payment flow.

Where AI written code is already good: test scaffolding, serialization boilerplate, CRUD handlers, config parsing, migration scripts, glue code between two libraries, regex you would have to look up anyway, and the boring 80% of a feature that nobody enjoys writing. I would merge most of that after a skim. The model has seen ten thousand examples of it and the correctness bar is low.

Where it is not good, and where the money and the incidents live: authentication and session handling, authorization checks, anything that moves money, concurrency and locking, retry and idempotency logic, input validation at trust boundaries, cryptography, and cache invalidation. These are the places where a plausible-looking implementation is wrong in a way that passes tests, passes review, and fails in production at 3am. The pattern is consistent: the model produces code that looks like the median of what it has seen, and the median of production code is not correct code.

The failure modes

Hallucinated APIs and packages

This is the one that has a name now. Slopsquatting is the practice of registering a package name that a language model tends to invent, then waiting for developers to install it. The research on package hallucination rates varies by study and by model, but the direction is consistent: models invent plausible package names at a non-trivial rate, and attackers have noticed. If your dependency installation is not gated, an AI written import line is an attack surface.

The API version of this is subtler. A model will call a method that existed in version 2 of a library and was removed in version 4. It will pass arguments in the order from an older release. It will import from a path that was reorganized two years ago. None of this is malicious and all of it is a build break or, worse, a silent behavior change if the old signature still exists with different semantics.

Silent logic errors that pass tests

This is the failure mode that costs the most, because your test suite is the thing that is supposed to catch it. AI written code tends to be tested by AI written tests, and both share the same misunderstanding of the requirement. If the model thinks pagination is zero-indexed and you think it is one-indexed, the test it writes will assert the model's assumption and pass. You get green CI and an off-by-one that surfaces as a missing first row in a report.

The other shape is the boundary condition. Off-by-one in a loop, a missing null check on a field that is nullable in the schema but not in the sample data, a timezone assumption baked into a date comparison. These are not exotic. They are the standard content of production incidents, and AI written code produces them at a rate that is not obviously better than a junior developer working from the same spec.

Security

The categories are boring and old: SQL injection through string concatenation, missing input validation on a boundary, hardcoded secrets in a config file the model helpfully generated, path traversal in a file handler, SSRF in a URL fetcher, and authorization checks that exist on the happy path but not on the error path. None of these are new. What is new is the volume. A team that used to write 200 lines of security-sensitive code a week now writes 800, and the review capacity did not scale with it.

Hardcoded secrets deserve a specific mention because the model does it constantly in examples. It will write api_key = "sk-..." in a sample because that is what the training data looks like. If that sample gets copied into a real file and committed, you have a credential in your git history. Secret scanning in CI catches most of this, but only if it runs on every commit and not just on pull requests.

License contamination

Models are trained on public code, and public code has licenses. The legal consensus is still forming, but the practical risk is real: a model can reproduce a distinctive function from a copyleft-licensed project closely enough that a lawyer would want to talk about it. The mitigation is not to avoid AI written code. It is to have a policy about which licenses are acceptable in your dependency tree and to run a license scanner, because the same scanner that catches a GPL dependency will flag a suspiciously similar code block if you configure it to.

How to review AI written code

The single highest-value rule I would put in a team handbook is a diff-size limit. Never merge an AI-generated diff over 300 lines without a human reading every line. The number is arbitrary but the principle is not: review quality collapses as diff size grows, and AI generated diffs grow fast because generation is cheap. A 300-line limit forces the author to split the work into reviewable chunks, which is good for humans too.

Beyond size, the checklist that matters is short. Does this code cross a trust boundary, and if so, is the input validated on the correct side? Does every error path either handle the error or propagate it, and is there any path that silently swallows it? Does every dependency in the diff actually exist, at the version claimed, with the API used? Does the code touch auth, money, or concurrency, and if so, has a human who understands that domain read it line by line?

Tooling does the mechanical part. A linter catches style and a subset of bugs. A software composition analysis tool catches known-vulnerable dependencies and, with configuration, license issues. A secret scanner catches credentials. Static analysis catches injection patterns and some taint flows. Wire all of these into CI so they run on every pull request, and make the build fail on findings rather than warn. A warning that nobody reads is not a control.

AI reviewing AI is where teams get into trouble. It works for the mechanical checks: does this function have a docstring, is this error handled, does this match the style guide. It works less well for the judgment calls, because the reviewer model shares the same blind spots as the author model. If the author misunderstood the requirement, the reviewer will likely agree with the misunderstanding. Use AI review as a first pass that catches the obvious, and keep a human on the diffs that touch the categories above. If you want a comparison of the review tools themselves, our AI code review tools breakdown covers the field, and the AI-powered code review tools comparison goes deeper on the ones that run in CI.

Metrics to track

You cannot manage what you do not measure, and most teams measuring AI written code are measuring the wrong thing. Lines of code generated is a vanity metric. Acceptance rate in the editor is a proxy for developer convenience, not for quality. The metrics that matter are the ones that describe what happened after the code shipped.

Defect escape rate is the first. Of the bugs found in production, what share came from AI-assisted pull requests versus human-only ones? You need to tag your PRs to answer this, and the tagging has to be honest. Revert rate is the second: what share of AI-assisted merges get reverted within 30 days, compared to human-only merges? Review time per PR is the third, and it is the one that reveals whether your review process is actually absorbing the volume or just rubber-stamping it.

DORA metrics split by AI-assisted versus human PRs is the more sophisticated version. Deployment frequency, lead time for changes, change failure rate, and time to restore. If AI written code is speeding up lead time but doubling change failure rate, you have not shipped faster, you have moved the cost from development to operations. The Google DORA research has consistently found that AI adoption correlates with throughput gains and, in some cohorts, with instability, which is exactly the tradeoff you are trying to measure.

A 30-day measurement plan you can copy: week one, tag every PR as AI-assisted or human-only and start collecting revert rate and review time. Week two, add defect escape rate by joining production incidents back to the originating PR. Week three, split your DORA metrics by the same tag. Week four, review the numbers and set a threshold. If AI-assisted change failure rate is more than 1.5x the human rate, tighten the review rules before you expand AI usage. If it is comparable, you have evidence to expand.

Guardrails that actually hold

Policy first. There is a list of files AI may never touch, and it should be written down and enforced. Auth and session code. Payment and billing. Anything under a directory named crypto or secrets. Infrastructure as code that provisions production. Database migrations that alter existing columns. The list is short and the point is that it is explicit, so a new hire knows the boundary without asking.

Prompt-level constraints and repository instruction files are the second layer. A CONTRIBUTING file or a tool-specific instruction file that tells the model the house rules: use the project's error type, do not add dependencies without approval, do not write to the database outside the repository layer, follow the existing test patterns. This does not eliminate bad output but it shifts the distribution. The model that knows your conventions produces code that needs less correction.

CI gates are the third layer and the only one that is actually enforced. Coverage delta: fail the build if the diff reduces coverage. Dependency allowlist: fail if a new dependency is added that is not on the approved list, which catches slopsquatting before install. SAST: fail on high-severity findings. Secret scanning: fail on any match. These are not novel controls. They are the same controls you would want for human-written code, applied with the assumption that the volume is higher and the reviewer attention is lower.

Accountability and disclosure

The question that comes up in audits is who owns an AI bug. The answer that survives audit is the same as for any other bug: the human who merged the pull request. The model is a tool, the same way a compiler is a tool, and the person who signed off on the change owns the outcome. Teams that try to route accountability to the tool find that auditors do not accept it, and they are right not to.

Disclosure norms are still forming. Some teams tag AI-assisted PRs in the description, some use a commit trailer, some do not disclose at all. The practical argument for disclosure is that it makes the metrics above possible. You cannot measure AI-assisted defect rate if you do not know which PRs were AI-assisted. The practical argument against is that it can create a stigma that discourages honest reporting. I would tag for measurement and keep the tag out of performance reviews.

Attribution for open-source-heavy AI output is the harder problem. If a model reproduces a function from a GPL project, the license obligations may attach to your code. The mitigation is a license scanner in CI and a policy about which licenses are acceptable, applied to both dependencies and, where the scanner supports it, to code similarity. This is not a solved problem in 2026 and anyone who tells you it is has not read the lawsuits.

What this means for tool choice

The tool you use to generate AI written code matters less than the process around it, but it is not irrelevant. Models differ in how often they invent APIs, how well they follow repository conventions, and how much they cost per useful diff. The current pricing snapshot shows a wide spread: openai/o1-pro at $150 per million input tokens and $600 per million output tokens, anthropic/claude-opus-4.7-fast at $30 in and $150 out, openai/gpt-5.5-pro at $30 in and $180 out, and cheaper options like openai/gpt-5-pro at $15 in and $120 out. Batch tiers roughly halve the input cost on some of these.

The practical implication is that the expensive models are worth it for the diffs that touch the categories above, and the cheap ones are fine for boilerplate. Routing by risk is a reasonable policy: use the frontier model for auth and money, use the cheap model for tests and glue. If you are still choosing a generator, our best AI code generators for 2026 roundup and the AI code generator guide cover the field, and the AI tools for coding guide covers the surrounding workflow.

Beetlix is our own product, and where it fits honestly is the review and measurement layer rather than the generation layer. If you are already generating code with one of the models above and you need the tagging, the metrics, and the CI gates described in this article, that is the problem it addresses. It is not a replacement for the human who owns the merge.

How this review was researched

This article is based on vendor documentation, the official pricing pages for the models named, public reporting from GitHub Octoverse and the Google DORA research program, and the live pricing data available at the time of writing. No testing was performed on any tool described here, and no benchmark numbers are claimed. Where a figure is cited, it comes from the source named in the sentence. Where a practice is recommended, it is a judgment call based on the failure modes documented in public incident reports and the security research on package hallucination.

What works

  • Covers the failure modes that actually cause incidents rather than the ones that make good demos
  • Gives concrete, enforceable rules: diff-size limits, CI gates, file-level policy
  • Separates the metrics that matter from vanity metrics like lines generated
  • Routes model spend by risk instead of treating all diffs the same
  • Names the accountability question directly instead of avoiding it

What doesn't

  • No single tool recommendation, which will frustrate readers looking for a purchase decision
  • The 300-line diff rule is a heuristic, not a measured threshold
  • License contamination section is honest about being unsolved, which is correct but unsatisfying
  • Assumes a team with CI infrastructure; solo developers get less value

The verdict

This is a process guide, not a tool review, and that is the right shape for the problem. The failure modes of AI written code are well documented and the mitigations are boring: smaller diffs, real CI gates, honest metrics, and a human who owns the merge. Teams that adopt those will get the throughput without the incidents. Teams that skip them will learn the same lessons from production.

FAQ

What is slopsquatting and why does it matter for AI written code?
Slopsquatting is registering a package name that a language model tends to invent, then waiting for developers to install it. Because models hallucinate plausible package names at a measurable rate, an AI written import line is a real attack surface. The mitigation is a dependency allowlist in CI that fails the build when a new package is added without approval.
How large an AI-generated diff is safe to merge without line-by-line review?
There is no measured safe threshold, but a 300-line limit is a reasonable heuristic. Review quality drops as diff size grows, and AI generated diffs grow fast because generation is cheap. A limit forces the author to split work into reviewable chunks, which helps human-written code too.
Who is accountable when AI written code causes a production incident?
The human who merged the pull request. The model is a tool, the same way a compiler is a tool, and auditors do not accept routing accountability to the tool. Teams that tag AI-assisted PRs for measurement should keep that tag out of performance reviews to avoid discouraging honest reporting.