For some reviews. DeepSeek V4 Pro had the best result of any model on CI/CD pull requests, and V4.1 Flash tied for first on SOC 2. On general code review Muse Spark 1.3 scored higher, and on security GLM 5.3 Flash scored higher for less.
DeepSeek code review: how DeepSeek V4 scores on real pull requests
IonWarp runs DeepSeek V4.1 Flash for Code Review and most other reviewers, on every plan, and benchmarks DeepSeek V4 Pro next to it. Both reviewed the same open-source pull requests as Muse Spark 1.3, GLM 5.3 Flash, GPT 5.6 Luna Pro and four review apps, and every finding was graded against known issues.[1]
DeepSeek leads one review outright: on 18 CI/CD pull requests, DeepSeek V4 Pro found 12 of 13 known pipeline defects, an F1 of 72.7%, for about 3 cents a review.[2] On general code review it trails: on 50 pull requests Muse Spark 1.3 scored 56.6%, DeepSeek V4 Pro 51.4% and V4.1 Flash 46.7%, at about 42, 16 and 12 cents a review.[1]
Where DeepSeek leads, and where it does not

CI/CD: DeepSeek V4 Pro leads
12 of 13 known pipeline defects found on 18 pull requests. V4.1 Flash found 11 for about a cent a review.[2]
CI/CD PerformanceOn Max

SOC 2: V4.1 Flash ties for first
13 of 15 known control gaps found on 20 pull requests, an F1 of 66.7%, tied with GLM 5.3 Flash, which took 34 seconds a review to V4.1 Flash's six minutes.[4]
SOC 2 ReviewOn Max

Code review: second and third of five models
Muse Spark 1.3 led the models. DeepSeek V4 Pro came second, and 70% of its findings matched a known issue; V4.1 Flash came third.[1]
Code ReviewOn Starter, Pro and Max

Security: GLM 5.3 Flash does better for less
On 13 security pull requests V4.1 Flash scored 62.5% at about 6 cents a review, and GLM 5.3 Flash 66.7% at about a cent. DeepSeek V4 Pro found only 4 of the 18 known issues.[5]
Security ReviewOn Starter, Pro and Max

Docs: Gemini 3.8 Flash leads
DeepSeek V4 Pro found the most stale docs, 15 of 17, but about half its findings matched no known issue. Gemini 3.8 Flash had the top F1, 68.6% to 65.2%.[6]
Docs FreshnessOn Pro and Max
How the benchmark grades a model
The same pull requests
Every model reviews the same public pull requests with the same reviewer instructions.
Known issues
Each pull request has an answer key. Code review uses the 50-PR code-review-bench from Martian;[7] the other reviews use smaller sets with provisional labels.
A grader matches each finding
A model grader, Claude Opus 4.5 on code review, matches each finding to a known issue. Precision is the share of findings that match one, recall the share of known issues found, and F1 combines the two.
Starter is free for 3 seats. Pro is $49 a month with 5 seats, and Max is $149 a month with 10 seats. Compare plans
Frequently asked questions
V4 Pro scored higher on our code benchmark (F1 51.4% to 46.7%) with fewer findings that matched no known issue, at about 16 cents a review against 12. IonWarp still runs V4.1 Flash for Code Review: on our production reviews it raised more serious findings per review at about a quarter of V4 Pro's cost.
We have not benchmarked Claude models as reviewers. Claude Opus 4.5 is the grader on the code review benchmark: it decides whether each finding matches a known issue.
In model fees on our benchmarks, V4.1 Flash cost about 12 cents a code review and 1 to 9 cents on the smaller specialist reviews; V4 Pro cost about 16 cents a code review. IonWarp plans include the model cost: Pro is $49/mo for the workspace.
Sources
- The code review benchmark: 50 pull requests, 158 known issues (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
- The CI/CD benchmark: 18 pull requests, 13 known pipeline defects (release of 2026-09-20) — IonWarp benchmark, read 2026-10-03.
- Pricing — CodeRabbit, read 2026-10-02.
- The SOC 2 review benchmark: 20 pull requests, 15 known control gaps (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
- The security review benchmark: 13 pull requests, 18 known issues (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
- The docs review benchmark: 19 pull requests, 17 known issues (release of 2026-09-22) — IonWarp benchmark, read 2026-10-03.
- withmartian/code-review-benchmark — Martian on GitHub, read 2026-10-03.
Get a review on your next pull request
Install IonWarp on GitHub. Starter is free for 3 people, with 15,000 credits to start.
Try for free