GuidesAI Code Review Tools: What the Diff Cannot Contain
76 modules, 169 import edges. A third of them can be reviewed from the diff alone. Nine of them cannot be reviewed from a diff at all.
CATEGORY · 16 ARTICLES
Guides76 modules, 169 import edges. A third of them can be reviewed from the diff alone. Nine of them cannot be reviewed from a diff at all.
GuidesTen seeded faults, every one a failure this project actually hit. Three produce no message for any debugger to read.
GuidesWe counted every turn in the sessions that built this site. The median run between two human sentences is 30 tool calls; the longest is 236.
GuidesSWE-bench Verified counted from the dataset rather than the paper. Every gold patch but one edits Python, which makes the leaderboard a Python leaderboard and almost nothing else.
GuidesOne question, one model, two configurations. Removing the harness cost the right answer and 72% of the loaded context, and the loop sent the transcript three times over.
GuidesNine files from one repository, measured by differencing two headless turns. The folklore constant undercounts every one of them, and the session was 31,979 tokens deep before anyone typed.
GuidesThe filter that builds SWE-bench keeps roughly 2.5% of the pull requests it sees. Point it at your own repository and you learn more than the leaderboard will tell you.
GuidesZero data retention, content exclusion, SOC 2. Every one of these controls has a documented gap, and the gap is in the vendor's own docs. Here's what to read before you buy seats.