Skip to content
IonWarpRouterTry for free

Best model for code review: benchmark (Sep 2026)

Muse Spark 1.3 has the highest F1 of the 5 models graded on every PR (56.6%) and is also the cheapest within 2 points of it, at $0.42 a review.

Reviewer / modelF1PrecisionRecall$ per reviewp50 timePRs
IonWarp Code · Muse Spark 1.356.6%65.3%50.0%$0.423m 18s50/50
Qodo v258.5%55.4%62.0%~$1.675m 40s50/50
IonWarp Code · DeepSeek V4 Pro51.4%70.3%40.5%$0.165m 59s50/50
Greptile v4.151.1%50.9%51.3%~$1.50–2.004m 58s50/50
IonWarp Code · GLM 5.3 Flash42.8%50.0%37.3%$0.022m 15s50/50
Cursor Bugbot51.4%56.9%46.8%~$1.00–1.507m 52s50/50
IonWarp Code · GPT 5.6 Luna Pro41.6%73.0%29.1%$0.121m 47s50/50
IonWarp Code · DeepSeek V4.1 Flash46.7%46.0%47.5%$0.1212m 32s50/50
CodeRabbit42.2%32.6%59.5%~$1.507m 34s50/50

Rows graded on every PR come first, ranked by Reviewer Score: F1 counts 3×, review time 1× and cost 1×. App prices (~) are list-price estimates; model prices are measured inference cost per review. The PRs and their expected issues come from the independent Martian code-review-bench; IonWarp ran the evaluation, which Martian has not certified.

50 PRs · snapshot · How we measure