Skip to content

Add ClawBench (live-website everyday-task agent benchmark) - #73

Open
reacher-z wants to merge 1 commit into
benchflow-ai:mainfrom
reacher-z:add-clawbench
Open

Add ClawBench (live-website everyday-task agent benchmark)#73
reacher-z wants to merge 1 commit into
benchflow-ai:mainfrom
reacher-z:add-clawbench

Conversation

@reacher-z

Copy link
Copy Markdown

Adds ClawBench (arXiv 2604.08523, TIGER-AI-Lab) to the agent-specific evaluation section, in the live-web cluster next to WebVoyager / Online-Mind2Web / REAL.

ClawBench tests whether web/computer-use agents complete everyday online tasks (purchases, bookings, job applications, email) on live production websites; a Chrome-extension + CDP layer blocks only the final write request so runs are safe. 153 tasks / 144 sites / 15 categories; best model in the paper (Claude Sonnet 4.6) = 33.3%.

Disclosure: I'm an author.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants