Browser Agent Benchmark

What does it cost an agent to get a browser task right?

Coding agents run the same eight browser tasks with different tools: the tasks from Stagehand's Playwright MCP token study, with the same model, harness and prompts for every tool. We measure tokens, cost, time and whether the answer was right. Every run, prompt and check is public.

Latest version of each tool

Medians per task run. Tokens include cache reads and writes. Lower is better, except success.

By package version

Every run of a tool version with this model counts toward its row, across workflow runs. Select a version for per-task results and the runs behind it.

Runs