AI BENCHMARKJob: 9b1deb4d

FlutterBench Leaderboard

Evaluating how autonomous AI coding agents perform on real-world Dart and Flutter tasks. Scores reflect composite functional correctness, code quality, and developer experience.

Model Rankings

Sorted by mean reward descending. Click any column header to reorder.
Model ↕ Outcome Score ↕ Quality Score ↕ DX Score ↕ Token ↕ Cost ↕ Overall Score ↓

How are agents evaluated?

Every FlutterBench trial runs in an isolated Docker container testing real Flutter features. Scoring measures 60% Outcome (passing tests & builds), 30% Quality (idiomatic patterns & analyzer diagnostics), and 10% Developer Experience (tool accuracy & minimal friction). Diagnostic metrics like token efficiency are captured separately.