Best model for API performance code review: benchmark (Sep 2026)
Muse Spark 1.3 has the highest F1 of the 6 models graded on every PR (56.0%) and is also the cheapest within 2 points of it, at $0.06 a review.
| Reviewer / model | F1 | Precision | Recall | $ per review | p50 time | PRs |
|---|---|---|---|---|---|---|
| IonWarp API Performance · Muse Spark 1.3 | 56.0% | 58.3% | 53.8% | $0.06 | 42 sec | 18/18 |
| IonWarp API Performance · GLM 5.3 Flash | 48.3% | 43.8% | 53.8% | <$0.01 | 49 sec | 18/18 |
| IonWarp API Performance · GPT 5.6 Luna Pro | 50.0% | 54.5% | 46.2% | $0.03 | 28 sec | 18/18 |
| IonWarp API Performance · DeepSeek V4 Pro | 52.2% | 60.0% | 46.2% | $0.04 | 1m 21s | 18/18 |
| IonWarp API Performance · DeepSeek V4.1 Flash | 53.3% | 47.1% | 61.5% | $0.04 | 6m 27s | 18/18 |
| IonWarp API Performance · Gemini 3.8 Flash | 41.7% | 45.5% | 38.5% | $0.07 | 47 sec | 18/18 |
| Greptile v4.1 | 76.9% | 83.3% | 71.4% | ~$1.50–2.00 | 4m 59s | 12/18 |
| Qodo v2 | 66.7% | 54.5% | 85.7% | ~$1.67 | 5m 43s | 12/18 |
| Cursor Bugbot | 71.4% | 71.4% | 71.4% | ~$1.00–1.50 | 7m 57s | 12/18 |
| CodeRabbit | 44.4% | 100.0% | 28.6% | ~$1.50 | 6m 0s | 12/18 |
Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. Provisional labels: IonWarp chose which issues count as API Performance Review findings, and no independent party has checked that choice.
18 PRs · snapshot · How we measure