Android Bench 2.0
AI-assisted software engineering has seen the emergence of several benchmarks to measure the capabilities of LLMs. Android developers face specific challenges that aren't covered by existing benchmarks, so we created one that focuses on a north star of high quality Android development.
Long-horizon task results
| Model - Agent | Pass rate (%) Average percentage of 30 tasks successfully resolved across all runs for each model |
arrow_range
Cl range (%)
Expected performance range, reflecting the results' statistical reliability (p-value < 0.05)
|
Completion rate (%) How close each run got to a complete solution, even when the task failed | Avg latency (h) Average time taken to solve 30 tasks across all runs | Avg cost ($) Average cost per full benchmark run |
|---|---|---|---|---|---|
|
GPT 6 Astra
chevron_right
codex |
28.0 | 13.3 — 42.0 | 82.2 | 7.9 | $375.7 |
|
Claude Fable 5 1
|