Best model for log & errors code review: benchmark (Sep 2026)
Gemini 3.8 Flash has the highest F1 of the 5 models graded on every PR (67.7%) and is also the cheapest within 2 points of it, at $0.10 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| IonWarp Log & Errors · Gemini 3.8 Flash | 67.7% | 67.7% | 67.7% | $0.10 | 1m 28s | 17/17 |
| Qodo v2 | 72.7% | 68.6% | 77.4% | ~$1.67 | 5m 52s | 17/17 |
| IonWarp Log & Errors · DeepSeek V4.1 Flash | 59.0% | 60.0% | 58.1% | $0.04 | 4m 8s | 17/17 |
| Greptile v4.1 | 67.8% | 71.4% | 64.5% | ~$1.50–2.00 | 5m 2s | 17/17 |
| Cursor Bugbot | 65.6% | 63.6% | 67.7% | ~$1.00–1.50 | 7m 34s | 17/17 |
| IonWarp Log & Errors · Muse Spark 1.3 | 43.6% | 50.0% | 38.7% | $0.08 | 56 sec | 17/17 |
| CodeRabbit | 59.2% | 52.5% | 67.7% | ~$1.50 | 8m 47s | 17/17 |
| IonWarp Log & Errors · GLM 5.3 Flash | 36.1% | 36.7% | 35.5% | $0.01 | 1m 14s | 17/17 |
| IonWarp Log & Errors · GPT 5.6 Luna Pro | 36.0% | 47.4% | 29.0% | $0.05 | 44 sec | 17/17 |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. Provisional labels: IonWarp chose which issues count as Log & Errors Review findings, and no independent party has checked that choice.
17 PRs · snapshot · How we measure