Best model for AI architecture code review: benchmark (Sep 2026)
DeepSeek V4.1 Flash has the highest F1 of the 8 models graded on every PR (47.1%) and is also the cheapest within 2 points of it, at $0.05 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| CodeRabbit | 72.7% | 80.0% | 66.7% | ~$1.50 | 7m 43s | 11/11 |
| IonWarp AI Architecture · GLM 5.3 | 47.1% | 36.4% | 66.7% | $0.07 | 1m 29s | 11/11 |
| IonWarp AI Architecture · GPT 5.6 Luna Pro | 35.3% | 27.3% | 50.0% | $0.05 | 51 sec | 11/11 |
| IonWarp AI Architecture · DeepSeek V4 Pro | 44.4% | 33.3% | 66.7% | $0.10 | 1m 52s | 11/11 |
| Greptile v4.1 | 60.0% | 75.0% | 50.0% | ~$1.50–2.00 | 5m 6s | 11/11 |
| IonWarp AI Architecture · DeepSeek V4.1 Flash | 47.1% | 36.4% | 66.7% | $0.05 | 6m 14s | 11/11 |
| IonWarp AI Architecture · GLM 5.3 Flash | 15.4% | 14.3% | 16.7% | $0.01 | 1m 2s | 11/11 |
| Cursor Bugbot | 50.0% | 40.0% | 66.7% | ~$1.00–1.50 | 7m 46s | 11/11 |
| IonWarp AI Architecture · Grok 4.6 | 34.8% | 23.5% | 66.7% | $0.28 | 3m 29s | 11/11 |
| IonWarp AI Architecture · Muse Spark 1.3 | 28.6% | 20.0% | 50.0% | $0.19 | 2m 4s | 11/11 |
| IonWarp AI Architecture · Gemini 3.8 Flash | 33.3% | 25.0% | 50.0% | $0.24 | 3m 50s | 11/11 |
| Qodo v2 | 25.0% | 50.0% | 16.7% | ~$1.67 | 5m 52s | 11/11 |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. Provisional labels: IonWarp chose which issues count as AI Architecture Review findings, and no independent party has checked that choice.
11 PRs · snapshot · How we measure