On IonWarp's benchmark of 50 pull requests and 158 known issues, Muse Spark 1.3 has the top Score (73.1; accuracy counts three times, speed and cost once), ahead of DeepSeek v4 Pro (52.6), GLM 5.3 Flash (41.6), GPT 5.6 Luna Pro (31.8) and DeepSeek v4.1 Flash (29.7), and of the Qodo, Greptile, Cursor Bugbot and CodeRabbit apps. On accuracy alone it was second, just behind Qodo.
The best LLM for code review, measured
We ran five models and four review apps on the same 50 open-source pull requests, which hold 158 known issues, and Claude Opus 4.5 graded every finding. On the benchmark's Score, which counts accuracy three times and speed and cost once, Muse Spark 1.3 came first with 73.1, ahead of Qodo (68.4), DeepSeek v4 Pro (52.6), Greptile (43.2), Cursor Bugbot (40.9) and CodeRabbit (7.7).[1] It found 79 of the 158 issues with 42 false alarms, in about three minutes a review.
On accuracy alone the Qodo app was a little ahead (F1 0.59 to 0.57). GPT 5.6 Luna Pro was the fastest model, under two minutes a review, and the most precise: 73% of its findings were real, but it found only 46 issues.[1] No model wins every column, or every kind of review: the table lists the leader for each.
What IonWarp runs, and why

Code Review, on every plan
Runs DeepSeek v4.1 Flash, picked on 1,958 production reviews of our own repos: it raised more P0 and P1 findings than DeepSeek v4 Pro (14.8 vs 13.5 per 100 reviews) at about a quarter of the cost.
Code ReviewOn Starter, Pro and Max

When a review is retried
A Code Review run that fails gets one retry, which can move to Muse Spark 1.3, the top Score above.
Code ReviewOn Starter, Pro and Max
The specialist reviewers
Security, docs, logs, performance, CI/CD, cost and SOC 2 are measured on 11 to 31 pull requests each, with provisional answer keys. Their leaders in the table are a guide; the models those reviewers run were picked on production reviews.
What our reviewers caught that all four apps missed

Hung processes that were never killed
Sentry started its flusher processes with the spawn context, so they were not multiprocessing.Process instances and the new isinstance check was always false.[3] DeepSeek v4.1 Flash flagged it; Qodo, Greptile, Bugbot and CodeRabbit did not.
Code ReviewOn Starter, Pro and Max

A migration that could break existing embeds
A Discourse migration wrote hosts with raw SQL and skipped the model's normalization, so hosts saved with http:// or a path might never match the new lookup.[4] DeepSeek v4.1 Flash flagged it.
Code ReviewOn Starter, Pro and Max

Clickjacking protection switched off
A Discourse change set X-Frame-Options to ALLOWALL and trusted the Referer header, which any client can fake.[5] DeepSeek v4.1 Flash flagged it.
Security ReviewOn Starter, Pro and Max

An error message for the wrong action
In Cal.com, the endpoint that turns off two-factor login answered with an error about backup code login.[6] Gemini 3.8 Flash flagged it.
Docs FreshnessOn Pro and Max
Starter is free for 3 seats. Pro is $49 a month with 5 seats, and Max is $149 a month with 10 seats. Compare plans
Frequently asked questions
We have not benchmarked Claude models as reviewers. Claude Opus 4.5 is the benchmark's grader: it decides whether each finding matches a known issue.
GPT 5.6 Luna Pro, at a median of 107 seconds a review, against 198 for Muse Spark 1.3 and 359 for DeepSeek v4 Pro. It was also the most precise, and it found the fewest issues.
Yes. Every reviewer has a pinned model on each plan, and you can switch any reviewer's model in settings. On Enterprise you can route reviews through your own AI gateway and model keys.
On IonWarp's benchmark pages: every model and app with its precision, recall, cost and time, and the graded findings on each pull request.
Sources
- The code review benchmark: 50 pull requests, 158 known issues (release of 2026-09-24) — IonWarp benchmark, read 2026-10-03.
- Pricing — CodeRabbit, read 2026-10-02.
- getsentry/sentry #93824, graded findings — IonWarp benchmark, read 2026-10-03.
- discourse #10 (embeddable hosts), graded findings — IonWarp benchmark, read 2026-10-03.
- discourse #4 (embed URL handling), graded findings — IonWarp benchmark, read 2026-10-03.
- calcom/cal.com #10600, graded findings — IonWarp benchmark, read 2026-10-03.
Get a review on your next pull request
Install IonWarp on GitHub. Starter is free for 3 people, with 15,000 credits to start.
Try for free