A real coding agent reads your docs and tries one onboarding task in a fresh sandbox. An independent judge verifies the evidence. You get a verdict and, on graded runs, an agent experience score.
Sorted by experience score. Experience measures friction in the tested task. Completion says whether it succeeded.
| # | Product | ScoreExperience score | ||
|---|---|---|---|---|
| Loading benchmark results... | ||||