WebMCPify Bench · Coming beta

Compare agent tasks with evidence.

A planned benchmark for measuring whether agents can discover, invoke and verify real website capabilities—across repeatable scenarios and explicit safety boundaries.

NOT AVAILABLE YET Scenarios and scoring are in development. No benchmark scores are published or implied.

What Bench will measure

Success is more than a green tool response.

Bench is designed to compare end-to-end task outcomes with the evidence needed to reproduce and challenge them.

01

Task completion

Did the intended application state actually change?

02

Discovery quality

Did the agent find the relevant capability and understand its schema?

03

Safety behavior

Were authentication, approval and risk boundaries respected?

04

Verification strength

Was success checked independently of the agent’s own report?

05

Reliability

Can the result be repeated across runs and environments?

06

Efficiency

How much time and agent work did a verified outcome require?

Request beta access

Help shape the first scenarios.

Leave one email. We'll only use it for this beta.