15 tools, run head-to-head. We score whether an agent can discover and choose the tool — then we run a live coding agent and score whether it can actually build. Every point traces to a check or a measured signal.
Discoverable — found, read & recommendedUsable — capability surface + agent builds with it