Task completion
Did the intended application state actually change?
WebMCPify Bench · Coming beta
A planned benchmark for measuring whether agents can discover, invoke and verify real website capabilities—across repeatable scenarios and explicit safety boundaries.
NOT AVAILABLE YET Scenarios and scoring are in development. No benchmark scores are published or implied.
What Bench will measure
Bench is designed to compare end-to-end task outcomes with the evidence needed to reproduce and challenge them.
Did the intended application state actually change?
Did the agent find the relevant capability and understand its schema?
Were authentication, approval and risk boundaries respected?
Was success checked independently of the agent’s own report?
Can the result be repeated across runs and environments?
How much time and agent work did a verified outcome require?