Best model for cost code review: benchmark (Sep 2026)
DeepSeek V4.1 Flash has the highest F1 of the 3 models graded on every PR (57.1%, $0.07 a review); GLM 5.3 Flash is the cheapest within 2 points of it, at $0.01 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| IonWarp Cost · GLM 5.3 Flash | 55.2% | 53.3% | 57.1% | $0.01 | 2m 23s | 20/20 |
| IonWarp Cost · DeepSeek V4.1 Flash | 57.1% | 47.6% | 71.4% | $0.07 | 10m 39s | 20/20 |
| IonWarp Cost · GPT-6 Luna Pro | 50.0% | 60.0% | 42.9% | $0.02 | 27 sec | 20/20 |
| CodeRabbit | N/A | N/A | N/A | ~$1.50 | N/A | N/A |
| Cursor Bugbot | N/A | N/A | N/A | ~$1.00–1.50 | N/A | N/A |
| Greptile v4.1 | N/A | N/A | N/A | ~$1.50–2.00 | N/A | N/A |
| Qodo v2 | N/A | N/A | N/A | ~$1.67 | N/A | N/A |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. Provisional labels: IonWarp chose which issues count as Cost Review findings, and no independent party has checked that choice.
20 PRs · snapshot · How we measure