Overview
- The benchmark, released Sept. 17, 2026, moves from short fixes to multi-day tasks such as building apps, adding major features, and converting cross-platform code to Android.
- Android Bench 2.0 uses a continuous completion rate that scores functionality, visual fidelity, regressions, and penalties for structural or instruction deviations.
- Google ran agent evaluations that pair models with developer harnesses, and the company found harness design materially changes developer outcomes and token efficiency.
- Early results show much lower top scores than before: GPT-6 Astra leads at about a 28% completion rate while frontier models top out at roughly 80% on cross-platform porting tasks.
- The tests show strengths on deterministic transforms and new-code generation but consistent failures on refactors, migrations, runtime validation, and tasks that require architectural understanding, a gap that will shape tooling and developer workflows.