Why coding agents produce bad code, big bills, and unreviewed merges

Most teams answer bad agent output with a longer instruction file. Longer files measurably make it worse. Here's what to cut, what to move into permissions, and what to enforce at the merge.

Oct 2, 2026 • 10 Minute Read

Please set an alt value for this image...
  • Software Development
  • AI & Data
  • DevOps

The three-minute version

None of these is a better prompt. If your teams point coding agents at production repositories, here are three things they can do to buy back velocity and code quality together.

  • Cut the instruction file down, then leave it alone. Every repository gets one file, CLAUDE.md for Claude Code or .github/copilot-instructions.md for Copilot, holding only what the agent can't work out for itself.
  • Scope every subagent explicitly. An agent spawned without a tool list inherits everything the parent model owns, and that's how weekend incidents start.
  • Gate the merge, then measure what clears the gate. Name the acceptance test in the spec before an agent touches the branch, and bind that test to a required status check. A specification nothing can enforce doesn't govern anything.

Intent tells the agent what to try. Permissions set the boundary, and verification decides whether the work gets merged. Most engineering organizations I talk to have written the first one, purchased the second one, and skipped the third. Start with the third.

Three layers govern a coding agent, and only the bottom two can stop it. Most teams have written the top one and skipped the bottom one.

An example: How your team’s long CLAUDE.md files can make output worse

Priya Sharma is a platform lead who runs a fourteen-person team at Globomantics, a fictional company. Last Tuesday, a pull request that cleared review on Friday broke checkout, because the generated code used a retry pattern her team abandoned two years ago and nobody wrote it down where an agent would find it. Her staff engineer proposed a shared prompt library. Priya killed it: "Nobody here is bad at prompting. We're bad at writing down what we already decided."

She's right, and the instruction file is where most teams then go wrong. It's context, prepended to every exchange, so a bloated one costs you on every turn. Researchers at ETH Zurich and LogicStar.ai tested repository context files across coding agents and found they raised inference cost more than 20 percent on average without improving task success. Anthropic's own documentation is direct about it: target under 200 lines, because "longer files consume more context and reduce adherence."

So your team should put every line through this test. Could the agent work it out alone, and if it could, would working it out cost more than the line costs you every turn? A dependency list fails both, because the agent reads the manifest in one call and your copy is cruft by Friday. A fifteen-line map of which directories are frozen and which conventions differ from the tool's defaults passes easily, because it encodes judgment rather than contents. Priya's file runs forty lines.

This part is genuinely hard. The line between judgment and contents isn't obvious until your team has cut the same file twice. My colleague Axel Sirota mapped the full maturity ladder from raw vibe coding up to versioned constitutional constraints. This piece is about what happens after the spec exists.

The same repository, two instruction files. Everything struck out on the left, the agent can read for itself in one call.

Why scoping subagents is important: They inherit more than you think

Notice what even a good instruction file can't do. Claude Code's documentation says it plainly: Claude "treats them as context, not enforced configuration. To block an action regardless of what Claude decides, use a PreToolUse hook instead." An instruction file asks, and a hook refuses.

That stops being academic the moment you delegate. A subagent an engineer spins up to tidy logging will happily refactor files under infra/, because an unscoped agent's boundary defaults to everything you own. In July 2025 a Replit agent deleted SaaStr founder Jason Lemkin's production database during a declared code freeze, then reported that rollback was impossible. It wasn't. Replit's fix wasn't a better prompt. They shipped dev and production separation, a planning-only mode, and one-click restore.

Make it standard practice to write the boundary into the instruction file so the agent understands your team’s intent, then enforce it in settings.json deny rules and pre-tool hooks so it can't cross the line. Require an explicit tool list for every subagent instead of an inherited one, and put CODEOWNERS on the instruction files so nobody edits the governance quietly.

An instruction file would have asked the agent to stay out of infra. The hook made staying out the only option.

Preventing AI quality issues with merge gating, and four numbers worth watching

Here's the layer almost nobody has built. A specification is only binding if something in your pipeline can refuse the work. The named acceptance test becomes a required status check on the protected branch, and code scanning runs on agent-authored pull requests exactly as it does on human ones. I've spent most of 2026 building GitHub Advanced Security courses around code scanning and dependency review, and that work convinced me of something: teams are bolting agents onto pipelines that were already too permissive, and the agent is only making the gap legible.

You might be thinking your team already does code review, so this is handled. Check the numbers first. Track revert rate on AI-authored pull requests, review cycles per pull request by author type, defect escape rate by author type, and time from open to first approval. If AI-authored changes clear review faster and revert more often, your reviewers are rubber-stamping volume they can't absorb, and no prompt fixes that.

GitClear's analysis of 623 million changed lines found moved code, the refactoring signal, fell from 13 percent of changed lines in 2023 to 3.8 percent in 2026, while block duplication climbed 81 percent. Apiiro's telemetry across tens of thousands of Fortune 50 repositories found AI-assisted code carried 76 percent fewer syntax errors and 322 percent more privilege escalation paths. Fewer syntax errors are nice, but syntax is exactly the kind of problem your linter is built to catch. Privilege escalation is the kind of failure that needs a specification and a gate.

The pattern above is what matters: when AI-authored changes clear review faster and revert more often, review has become a formality.

Your token bill is a missing termination condition

Spiraling spend gets treated as a pricing problem, and it's a control problem. Since June 1, 2026, Copilot meters in GitHub AI Credits at one cent per credit, drawn against input, output, and cached tokens at published API rates. An agent with a vague goal and no acceptance test compensates by exploring, and you buy the exploration by the token. Give it a named test to satisfy and the loop has somewhere to stop.

None of this survives an untrained team

Priya can write a perfect instruction file this afternoon and it changes nothing if the other thirteen people keep vibe coding. Anthropic's randomized trial of 52 mostly junior developers found the AI-assisted group scored 17 percentage points lower on a comprehension quiz, 50 percent against 67 percent. The ones who asked conceptual questions held their mastery.

What changes the outcome is how people use the tool. You can teach that, and people need somewhere cheap to get it wrong. I work for Pluralsight, and Pluralsight sells training, so weigh this accordingly. For engineering organizations, AI Ready runs AI-assisted coding through agent development to agentic orchestration, with hands-on labs, pre-configured sandboxes, and a capstone on real codebases where being wrong costs a reset instead of an outage. When the gap is organization-wide instead, AI Academy covers that breadth.

Capping the instruction file is my call and not a vendor's, and it's the recommendation most likely to age badly as context gets cheaper. I'd still cap it today, because adherence and cost both degrade today. Feynman's first principle was that you mustn't fool yourself, and you are the easiest person to fool. AI coding tools are remarkably good at helping engineers fool themselves about velocity. Write the intent down. Enforce the boundary in configuration. Then make the merge prove itself. Priya gets a Tuesday with one ticket on it instead of three.

Keep reading, from colleagues who got here first

I'm standing on a lot of Pluralsight work in this piece, and the originals are better than my compression of them.

Tim Warner

Tim W.

Timothy Warner is a Microsoft Most Valuable Professional (MVP) in Cloud and Datacenter Management who is based in Nashville, TN. His professional specialties include Microsoft Azure, cross-platform PowerShell, and all things Windows Server-related. You can reach Tim via Twitter (@TechTrainerTim), LinkedIn or his blog, AzureDepot.com.

More about this author