Skip to content
IonWarpRouterTry for free
Guide

The best LLM for code review, measured

We ran five models and four review apps on the same 50 open-source pull requests, which hold 158 known issues, and Claude Opus 4.5 graded every finding. On the benchmark's Score, which counts accuracy three times and speed and cost once, Muse Spark 1.3 came first with 73.1, ahead of Qodo (68.4), DeepSeek v4 Pro (52.6), Greptile (43.2), Cursor Bugbot (40.9) and CodeRabbit (7.7).[1] It found 79 of the 158 issues with 42 false alarms, in about three minutes a review.

On accuracy alone the Qodo app was a little ahead (F1 0.59 to 0.57). GPT 5.6 Luna Pro was the fastest model, under two minutes a review, and the most precise: 73% of its findings were real, but it found only 46 issues.[1] No model wins every column, or every kind of review: the table lists the leader for each.

What IonWarp runs, and why

  • Code Review, on every plan

    Runs DeepSeek v4.1 Flash, picked on 1,958 production reviews of our own repos: it raised more P0 and P1 findings than DeepSeek v4 Pro (14.8 vs 13.5 per 100 reviews) at about a quarter of the cost.

    Code ReviewOn Starter, Pro and Max

  • When a review is retried

    A Code Review run that fails gets one retry, which can move to Muse Spark 1.3, the top Score above.

    Code ReviewOn Starter, Pro and Max

  • The specialist reviewers

    Security, docs, logs, performance, CI/CD, cost and SOC 2 are measured on 11 to 31 pull requests each, with provisional answer keys. Their leaders in the table are a guide; the models those reviewers run were picked on production reviews.

What our reviewers caught that all four apps missed

  1. Hung processes that were never killed

    Sentry started its flusher processes with the spawn context, so they were not multiprocessing.Process instances and the new isinstance check was always false.[3] DeepSeek v4.1 Flash flagged it; Qodo, Greptile, Bugbot and CodeRabbit did not.

    Code ReviewOn Starter, Pro and Max

  2. A migration that could break existing embeds

    A Discourse migration wrote hosts with raw SQL and skipped the model's normalization, so hosts saved with http:// or a path might never match the new lookup.[4] DeepSeek v4.1 Flash flagged it.

    Code ReviewOn Starter, Pro and Max

  3. Clickjacking protection switched off

    A Discourse change set X-Frame-Options to ALLOWALL and trusted the Referer header, which any client can fake.[5] DeepSeek v4.1 Flash flagged it.

    Security ReviewOn Starter, Pro and Max

  4. An error message for the wrong action

    In Cal.com, the endpoint that turns off two-factor login answered with an error about backup code login.[6] Gemini 3.8 Flash flagged it.

    Docs FreshnessOn Pro and Max

Starter is free for 3 seats. Pro is $49 a month with 5 seats, and Max is $149 a month with 10 seats. Compare plans

FAQ

Frequently asked questions

Sources

  1. The code review benchmark: 50 pull requests, 158 known issues (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
  2. Pricing — CodeRabbit, read 2026-10-02.
  3. getsentry/sentry #93824, graded findings — IonWarp benchmark, read 2026-10-03.
  4. discourse #10 (embeddable hosts), graded findings — IonWarp benchmark, read 2026-10-03.
  5. discourse #4 (embed URL handling), graded findings — IonWarp benchmark, read 2026-10-03.
  6. calcom/cal.com #10600, graded findings — IonWarp benchmark, read 2026-10-03.

Get a review on your next pull request

Install IonWarp on GitHub. Starter is free for 3 people, with 15,000 credits to start.

Try for free