Best model for CI/CD code review: benchmark (Sep 2026)
DeepSeek V4 Pro has the highest F1 of the 5 models graded on every PR (72.7%) and is also the cheapest within 2 points of it, at $0.03 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| IonWarp CI/CD Performance · DeepSeek V4 Pro | 72.7% | 60.0% | 92.3% | $0.03 | 1m 22s | 18/18 |
| IonWarp CI/CD Performance · GLM 5.3 Flash | 55.0% | 40.7% | 84.6% | <$0.01 | 50 sec | 18/18 |
| IonWarp CI/CD Performance · DeepSeek V4.1 Flash | 59.5% | 45.8% | 84.6% | $0.01 | 54 sec | 18/18 |
| IonWarp CI/CD Performance · GPT 5.6 Luna Pro | 54.5% | 45.0% | 69.2% | $0.01 | 33 sec | 18/18 |
| IonWarp CI/CD Performance · Muse Spark 1.3 | 44.4% | 42.9% | 46.2% | $0.03 | 49 sec | 18/18 |
| CodeRabbit | N/A | N/A | N/A | ~$1.50 | N/A | N/A |
| Cursor Bugbot | N/A | N/A | N/A | ~$1.00–1.50 | N/A | N/A |
| Greptile v4.1 | N/A | N/A | N/A | ~$1.50–2.00 | N/A | N/A |
| Qodo v2 | N/A | N/A | N/A | ~$1.67 | N/A | N/A |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. Provisional labels: IonWarp chose which issues count as CI/CD Performance findings, and no independent party has checked that choice.
18 PRs · snapshot · How we measure