Code review is now the slowest part of custom software delivery, and the teams commissioning that work are usually the last to figure it out. The agent writes the branch overnight. Tests run themselves.
Then the diff lands in a queue that still waits on a human senior engineer to read it carefully, ask the right questions, and decide whether it ships. That queue is where modern pipelines back up, and most of the assumptions buyers carry about agent-assisted engineering fall apart the moment they look at it.
The shift itself has been covered well. A recent Programming Insider feature argues that the pull request is becoming the new prompt — the unit of work an agent produces, submits, and defends. The myths worth clearing up are the ones that keep clients and teams from reorganizing around that reality.
Myth: The Agent Lives Inside the IDE
Most engineering leaders still picture a smarter autocomplete sitting next to a developer, suggesting the next line. That picture is a year out of date. The agent now runs on its own branch, picks up an issue from the backlog, drafts a plan, writes the code, runs the tests, and opens the pull request for a human to review.
GitHub's own documentation describes the Copilot coding agent working exactly this way: researching the repository, pushing commits to a draft PR, and waiting for a reviewer to approve before CI runs. The work happens outside the editor. The developer meets it at the diff.
Myth: More Code Shipped Means Faster Delivery
The appeal of agentic coding is throughput. Twenty PRs a week instead of five. The catch is that throughput of generated code is a different number than throughput of merged, reviewed, tested, and shipped code.
Review capacity has not changed. If anything it has shrunk, because reviewers are now reading unfamiliar code written by something that can't explain its own reasoning in a hallway conversation.
So the bottleneck moved. It used to sit at authorship. It now sits at review, CI, and test triage. A team that doubles its agent output without doubling its review bandwidth will watch its merge queue stretch, not shrink.
Myth: Review Is Just a Rubber Stamp Once Tests Pass
Agent-authored PRs pass tests all the time, and passing tests is a long way from being correct. Agents are good at producing code that compiles, satisfies the obvious test cases, and reads plausibly. They're weaker at the context around a change — the unwritten constraints a senior engineer carries in their head about why a module is structured the way it is.
Reviewers now have to look at a few specific things on agent diffs:
- Scope creep. Agents often fix the stated issue and then helpfully refactor three nearby files. Each unrelated change is a merge risk the ticket never asked for.
- Invented APIs. Function calls that look right but reference methods that don't exist, or exist in a different version of a library.
- Shallow tests. New tests that assert the code does what it does, rather than what the requirement says it should do.
- Silent assumptions. Config changes, migration scripts, or dependency bumps that work in isolation and break in staging.
Teams building a software partnership should ask how a vendor's reviewers handle these patterns, not whether they use agents at all. The guidance in Google's code review guide was written for human authors, and the principles translate — but the failure modes to look for are different.
Myth: Agents Replace Senior Engineers
The opposite is closer to the truth. Agents absorb work that used to belong to mid-level engineers — scaffolding, test fixtures, boilerplate CRUD, obvious bug fixes. What's left leans heavier on senior judgment: architecture, review, debugging edge cases, deciding what the agent should not be allowed to touch.
A team commissioning custom software today should expect a different staffing shape than a team buying the same project three years ago. Fewer junior hands typing. More senior hands reading.
The hourly rate on the invoice may not drop. The output per hour should.
Myth: The Buyer Doesn't Need to Care About the Pipeline
If you're the one commissioning the build, the merge pipeline used to be a black box. The vendor handled it. Agent-authored work changes that calculation, because the quality of what gets shipped now depends on review discipline you're paying for but can't see from the outside.
A few questions worth asking before the contract is signed:
- Who reviews agent PRs. Is it a senior engineer, a peer agent, or whoever is free? The answer tells you how seriously the vendor treats review as a bottleneck.
- What CI must pass before merge. Unit tests alone are not enough for agent-generated code. Ask about integration tests, security scanning, and dependency review.
- How scope is enforced. Does the team reject PRs that drift beyond the ticket, or merge them because the extra changes look harmless?
- What the human sign-off looks like. A reviewer who approves 40 PRs a day is not reading them. Ask what the realistic daily review load is per engineer.
The pull request is where custom software is governed now. The teams that treat review, CI, and testing as the real constraint — and staff them that way — are the ones whose agent-assisted builds actually ship on time. The rest will keep producing more code and merging less of it.



