AI Coding Tools: What Worked and What Did Not, Six Months In

Six months into the AI coding tool rollout, the picture has clarified enough to say what worked, what did not, and what the next six months should look like. The tool the developer has been using, the tool the developer…

Dark cinematic editorial image for AI Coding Tools: What Worked and What Did Not, Six Months In - abstract cyan and electric blue digital composition in deep black, hacker aesthetic, no text no logos

5 MIN READ

Picture the AI coding demo the board watched six months ago. Cursor, Copilot, Claude Code, the dev relished the autocomplete, the staff engineer ran a real refactor on stage, the engineering lead showed the cycle time number drop in real time. The board approved the rollout. Six months later, the demo has aged badly. Cycle time is down in some places, up in others. The senior engineer now spends Tuesday afternoons reviewing AI generated pull requests. The platform org quietly hired a reviewer whose entire job amounts to catching the regressions Cursor ships on the second pass. The honest framing is this: AI coding tools delivered exactly the productivity gain the demo promised, plus a tax the demo never mentioned, and the engineering org that treats the tax as a known cost is the one that comes out ahead.

Lineage worth knowing. The first six months of the AI coding rollout produced a track record. The boilerplate and the test got faster. The documentation stopped being skipped. The small refactor stopped being deferred. The review burden climbed, the architecture debt accumulated faster, and the security regression rate became the line the CISO office now tracks in the monthly report. Three months in, most engineering leads thought the gains were free. Six months in, the same leads are budgeting for the tax.

What actually shipped faster

The boilerplate and the test serve as the clearest win. A staff engineer who used to spend half a sprint writing scaffolding now spends the same half sprint shipping the feature, because Cursor and Claude Code produce the scaffolding in the time the engineer used to spend typing it. The generated test cases land alongside the code, and a meaningful share of the test cases the AI writes actually catch the bug they claim to. The documentation sits as the second gain, and it lands where the engineer used to skip. The README gets written at the same time the PR gets opened. The inline comment stops being a TODO. The onboarding doc for the new joiner stops being the postmortem the senior engineer writes after the new joiner has already quit. The small refactor runs as the third gain, and it counts as the one nobody celebrates because it shows up as fewer stale branches rather than more commits. An engineer proposes a rename, the AI produces the diff across 200 files, the engineer reviews the diff in 20 minutes, the PR lands the same afternoon. Multiply that across a 200 person engineering org and the time saved becomes a number the platform lead can defend.

What got slower or worse

The review burden sits as the cost nobody priced into the demo. A senior IC used to review a 200 line PR in 30 minutes. The same senior IC now reviews a 600 line AI generated PR in 90 minutes, and the 600 lines contain most of the right code plus one SQL injection the AI invented with full confidence. The code review process has not caught up. Architecture debt accumulates next, and the pattern repeats across every codebase the engineering org has shipped in the last six months. The AI ships a utility that duplicates a utility that already exists. The AI picks a pattern that contradicts the pattern the staff engineer established last quarter. The AI forgets the error handling the platform team built into the shared library. Each individual miss is small. The aggregate is the architecture drift the engineering lead now spends one cleanup sprint per quarter unwinding. The security regression rounds out the cost, and the CISO office has the data. Snyk’s 2025 survey of enterprise engineering orgs found a meaningful share of AI generated pull requests introduced at least one security defect that the manual review did not catch, with the most common offenders being hardcoded secrets, missing auth checks, and SQL injection that the AI produced with the same confidence it produced the rest of the change. The pattern Snyk reported matches what the platform teams at the larger engineering orgs are now seeing in their own dashboards.

What the next six months should fix

Tune the tool to the codebase. The Cursor rule file the platform team writes in week one becomes the rule file every engineer on the team uses for the next quarter. Custom rules, example patterns from the existing codebase, the test conventions the staff engineer established last year. A week of tuning saves a quarter of architecture debt.

Add the review automation. The static analysis, the secret scanner, the security linter, the architecture checker, all of it running on the AI generated diff before the senior IC reviews the code. A sprint of automation setup catches the regression the manual review would have missed, and the cost gets paid back inside the first quarter.

Measure the right thing. Lines of code shipped counts as the metric the board celebrated in the demo and the metric nobody should be reporting now. Cycle time, bug rate, security regression rate, and architecture debt per quarter run as the four numbers the engineering lead should put in the next retro. The retro that reports those four numbers honestly amounts to the retro that turns the AI coding tool into the productivity gain the demo promised.

Abstract six month retrospective as glowing cyan timeline with milestones on a dark navy surface, dramatic chiaroscuro lighting from above.
AI coding 6 month retro in 2026: 3 gains (boilerplate, documentation, small refactor), 3 costs (review burden, architecture debt, security regression), 3 moves for the next 6 months.

The bottom line

Tune the tool to the codebase, add the review automation, measure the right thing. The AI coding tool the engineering org rolled out six months ago runs as a real productivity gain, and the engineering lead who treats the review tax as a known cost becomes the one who keeps the gain when the board stops believing the demo.

Sources & Further Reading

All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.

Spotted an error? Email the editor. Corrections are issued with a visible correction note.

Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.

Continue reading