Best model for code review: benchmark (Sep 2026)
Muse Spark 1.3 has the highest F1 of the 5 models graded on every PR (56.6%) and is also the cheapest within 2 points of it, at $0.42 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| IonWarp Code · Muse Spark 1.3 | 56.6% | 65.3% | 50.0% | $0.42 | 3m 18s | 50/50 |
| Qodo v2 | 58.5% | 55.4% | 62.0% | ~$1.67 | 5m 40s | 50/50 |
| IonWarp Code · DeepSeek V4 Pro | 51.4% | 70.3% | 40.5% | $0.16 | 5m 59s | 50/50 |
| Greptile v4.1 | 51.1% | 50.9% | 51.3% | ~$1.50–2.00 | 4m 58s | 50/50 |
| IonWarp Code · GLM 5.3 Flash | 42.8% | 50.0% | 37.3% | $0.02 | 2m 15s | 50/50 |
| Cursor Bugbot | 51.4% | 56.9% | 46.8% | ~$1.00–1.50 | 7m 52s | 50/50 |
| IonWarp Code · GPT 5.6 Luna Pro | 41.6% | 73.0% | 29.1% | $0.12 | 1m 47s | 50/50 |
| IonWarp Code · DeepSeek V4.1 Flash | 46.7% | 46.0% | 47.5% | $0.12 | 12m 32s | 50/50 |
| CodeRabbit | 42.2% | 32.6% | 59.5% | ~$1.50 | 7m 34s | 50/50 |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. The PRs and their expected issues come from the independent Martian code-review-bench; IonWarp ran the evaluation, which Martian has not certified.
50 PRs · snapshot · How we measure