Best model for security code review: benchmark (Sep 2026)
GLM 5.3 Flash has the highest F1 of the 7 models graded on every PR (66.7%) and is also the cheapest within 2 points of it, at $0.01 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| IonWarp Security · GLM 5.3 Flash | 66.7% | 83.3% | 55.6% | $0.01 | 1m 38s | 13/13 |
| IonWarp Security · DeepSeek V4.1 Flash | 62.5% | 71.4% | 55.6% | $0.06 | 5m 56s | 13/13 |
| Qodo v2 | 73.3% | 91.7% | 61.1% | ~$1.67 | 7m 16s | 13/13 |
| IonWarp Security · Muse Spark 1.3 | 44.4% | 66.7% | 33.3% | $0.26 | 2m 30s | 13/13 |
| Greptile v4.1 | 53.8% | 87.5% | 38.9% | ~$1.50–2.00 | 5m 2s | 13/13 |
| IonWarp Security · Grok 4.6 | 53.3% | 66.7% | 44.4% | $0.58 | 7m 42s | 13/13 |
| IonWarp Security · Gemini 3.8 Flash | 51.6% | 61.5% | 44.4% | $0.52 | 7m 10s | 13/13 |
| CodeRabbit | 54.1% | 52.6% | 55.6% | ~$1.50 | 9m 26s | 13/13 |
| Cursor Bugbot | 50.0% | 100.0% | 33.3% | ~$1.00–1.50 | 8m 23s | 13/13 |
| IonWarp Security · DeepSeek V4 Pro | 34.8% | 80.0% | 22.2% | $0.13 | 2m 40s | 13/13 |
| IonWarp Security · GPT 5.6 Luna Pro | 28.6% | 100.0% | 16.7% | $0.11 | 1m 16s | 13/13 |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. Provisional labels: IonWarp chose which issues count as Security Review findings, and no independent party has checked that choice.
13 PRs · snapshot · How we measure