Overview / Description
APIEval-20 is an AI benchmark from KushoAI that measures how well API-testing agents can find bugs when given only a JSON schema and a single sample payload. It runs an agent's generated test suite against live reference APIs that contain planted bugs, then scores bug-detection accuracy, coverage, and efficiency objectively. The benchmark spans 20 scenarios across 7 domains, so results reflect performance on varied real-world API shapes rather than a single case. For developers and QA teams building or evaluating AI test-generation agents, APIEval-20 provides a repeatable yardstick to compare approaches and see where an agent misses edge cases. It's positioned as a research-grade benchmark for the API-testing-agent space rather than a tool you point at your own production APIs.
Used For
APIEval-20 is used by developers and QA teams to benchmark AI API-testing agents on bug detection, coverage, and efficiency using only a schema and a sample payload.
Pricing
Pros & Cons
Pros
• Scores agents from just a JSON schema and one sample payload • Tests against live reference APIs with planted bugs • 20 scenarios across 7 domains for varied coverage • Objective metrics: bug detection, coverage, efficiency
Cons
• A benchmark, not a tool to test your own production APIs • Aimed at developers building API-testing agents • Narrow, research-focused scope
Questions & Answers
Reviews & Ratings
0 reviews
Sign in to rate and review APIEval-20.
Sign in to reviewNo reviews yet. Be the first to review APIEval-20!