Skip to content
IonWarpRouterTry for free
Guide

DeepSeek code review: how DeepSeek V4 scores on real pull requests

IonWarp runs DeepSeek V4.1 Flash for Code Review and most other reviewers, on every plan, and benchmarks DeepSeek V4 Pro next to it. Both reviewed the same open-source pull requests as Muse Spark 1.3, GLM 5.3 Flash, GPT 5.6 Luna Pro and four review apps, and every finding was graded against known issues.[1]

DeepSeek leads one review outright: on 18 CI/CD pull requests, DeepSeek V4 Pro found 12 of 13 known pipeline defects, an F1 of 72.7%, for about 3 cents a review.[2] On general code review it trails: on 50 pull requests Muse Spark 1.3 scored 56.6%, DeepSeek V4 Pro 51.4% and V4.1 Flash 46.7%, at about 42, 16 and 12 cents a review.[1]

Where DeepSeek leads, and where it does not

  • CI/CD: DeepSeek V4 Pro leads

    12 of 13 known pipeline defects found on 18 pull requests. V4.1 Flash found 11 for about a cent a review.[2]

    CI/CD PerformanceOn Max

  • SOC 2: V4.1 Flash ties for first

    13 of 15 known control gaps found on 20 pull requests, an F1 of 66.7%, tied with GLM 5.3 Flash, which took 34 seconds a review to V4.1 Flash's six minutes.[4]

    SOC 2 ReviewOn Max

  • Code review: second and third of five models

    Muse Spark 1.3 led the models. DeepSeek V4 Pro came second, and 70% of its findings matched a known issue; V4.1 Flash came third.[1]

    Code ReviewOn Starter, Pro and Max

  • Security: GLM 5.3 Flash does better for less

    On 13 security pull requests V4.1 Flash scored 62.5% at about 6 cents a review, and GLM 5.3 Flash 66.7% at about a cent. DeepSeek V4 Pro found only 4 of the 18 known issues.[5]

    Security ReviewOn Starter, Pro and Max

  • Docs: Gemini 3.8 Flash leads

    DeepSeek V4 Pro found the most stale docs, 15 of 17, but about half its findings matched no known issue. Gemini 3.8 Flash had the top F1, 68.6% to 65.2%.[6]

    Docs FreshnessOn Pro and Max

How the benchmark grades a model

  1. The same pull requests

    Every model reviews the same public pull requests with the same reviewer instructions.

  2. Known issues

    Each pull request has an answer key. Code review uses the 50-PR code-review-bench from Martian;[7] the other reviews use smaller sets with provisional labels.

  3. A grader matches each finding

    A model grader, Claude Opus 4.5 on code review, matches each finding to a known issue. Precision is the share of findings that match one, recall the share of known issues found, and F1 combines the two.

Starter is free for 3 seats. Pro is $49 a month with 5 seats, and Max is $149 a month with 10 seats. Compare plans

FAQ

Frequently asked questions

Sources

  1. The code review benchmark: 50 pull requests, 158 known issues (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
  2. The CI/CD benchmark: 18 pull requests, 13 known pipeline defects (release of 2026-09-20) — IonWarp benchmark, read 2026-10-03.
  3. Pricing — CodeRabbit, read 2026-10-02.
  4. The SOC 2 review benchmark: 20 pull requests, 15 known control gaps (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
  5. The security review benchmark: 13 pull requests, 18 known issues (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
  6. The docs review benchmark: 19 pull requests, 17 known issues (release of 2026-09-22) — IonWarp benchmark, read 2026-10-03.
  7. withmartian/code-review-benchmark — Martian on GitHub, read 2026-10-03.

Get a review on your next pull request

Install IonWarp on GitHub. Starter is free for 3 people, with 15,000 credits to start.

Try for free