What does it cost an agent to get a browser task right?
Coding agents run the same eight browser tasks with different tools: the tasks from Stagehand's Playwright MCP token study, with the same model, harness and prompts for every tool. We measure tokens, cost, time and whether the answer was right. Every run, prompt and check is public.
Latest version of each tool
Medians per task run. Tokens include cache reads and writes. Lower is better, except success.
By package version
Every run of a tool version with this model counts toward its row, across workflow runs. Select a version for per-task results and the runs behind it.
Runs
No published benchmark runs yet.