JS Wei (Jack) Sun

1.02M-PR audit: AI code review cuts days per KLOC, adds 8% more smells

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality huggingface.co

Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a primarily human review process toward AI-supported review processes in which large language model (LLM) reviewers and AI agent reviewers participate alongside human reviewers. However, we still lack empirical evidence on how this transition affects review effic

References

GitClear 2026 AI Code Quality Research gitclear.com

AI accelerates code production [but] refactoring activity has dropped by 70% and code duplication has risen by 81% since 2023

daily.dev — ‘AI Code Review Limits: Why AI Reviewing AI Fails’ daily.dev

LLMs often fail to catch vulnerabilities in code they generated themselves — failing roughly 64.5% of the time — because they share the same training distribution blind spots as the generator

Wessel et al., EMSE 2022 (TU Eindhoven mirror) aserebre.win.tue.nl

bot adoption is associated with an increase in the number of monthly merged pull requests and a corresponding decrease in non-merged PRs [and] a significant reduction in human-to-human communication

SoftwareSeni — ‘Why Agent-Generated Code is Breaking the PR Review Model’ softwareseni.com

traditional PRs are viewed as a noisy, blocking formality… human oversight [should move] upstream — focusing on rigorous review of the initial specifications rather than line-by-line inspection of agent-generated artifacts

DeepSource — AI Code Review Benchmarks deepsource.com

SWRBench tested five LLM-based tools across 1,000 GitHub pull requests and found that four out of five had a precision rate below 10%

Medium (binbash) — ‘From Bugbot’s $40/month to Codebot’ medium.com

Bugbot transitioned to a $40/month premium tier, leading some developers to seek cheaper, custom-built alternatives using open-source frameworks like Claude Code

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare