Browse

APIEval-20

Updated June 30, 2026

Overview / Description

APIEval-20 is an AI benchmark from KushoAI that measures how well API-testing agents can find bugs when given only a JSON schema and a single sample payload. It runs an agent's generated test suite against live reference APIs that contain planted bugs, then scores bug-detection accuracy, coverage, and efficiency objectively. The benchmark spans 20 scenarios across 7 domains, so results reflect performance on varied real-world API shapes rather than a single case. For developers and QA teams building or evaluating AI test-generation agents, APIEval-20 provides a repeatable yardstick to compare approaches and see where an agent misses edge cases. It's positioned as a research-grade benchmark for the API-testing-agent space rather than a tool you point at your own production APIs.

Used For

APIEval-20 is used by developers and QA teams to benchmark AI API-testing agents on bug detection, coverage, and efficiency using only a schema and a sample payload.

Pricing

Free

$0/month

Free benchmark resource published by KushoAI.

View pricing

Pros & Cons

Pros

• Scores agents from just a JSON schema and one sample payload • Tests against live reference APIs with planted bugs • 20 scenarios across 7 domains for varied coverage • Objective metrics: bug detection, coverage, efficiency

Cons

• A benchmark, not a tool to test your own production APIs • Aimed at developers building API-testing agents • Narrow, research-focused scope

Questions & Answers

Reviews & Ratings

0 reviews

5
0%
4
0%
3
0%
2
0%
1
0%

Sign in to rate and review APIEval-20.

Sign in to review

No reviews yet. Be the first to review APIEval-20!

Try APIEval-20 free