Browse

Cekura Bench

Updated October 9, 2026

Overview / Description

Cekura Bench is an AI developer tool that benchmarks real-time speech-to-speech models by running them through actual phone calls, so teams can choose a voice model on evidence rather than vendor claims. Each model is tested on 82 real call scenarios, 59 appointment bookings and 23 Medicare intake calls, and every scenario is run three times to expose inconsistency. The results are scored across reliability (a pass-cubed metric that rewards models passing all three runs), success rate, data accuracy, stalled calls, response time, and cost per minute. What makes the benchmark usable rather than just a leaderboard is transparency: every call is published with its transcript, tool calls, and detailed scores, and you can run any two models head to head. A failure-analysis view shows exactly which scenarios break a given model and how different models fail differently, which is often more decisive than an aggregate score. For engineers and product managers asking what is the best AI for a production voice agent, Cekura Bench turns that question into comparable, inspectable data. The benchmark itself is free to browse; the per-minute figures shown are each model provider's own pricing, not Cekura's.

Used For

Comparing and selecting real-time speech-to-speech AI models for voice agents using transparent, call-based benchmark data.

Pricing

Plan

Free

Free to browse the benchmark; per-minute rates shown are each model provider's own pricing

View pricing

GPT Realtime 2.1

$0.08/month

$0.085 per minute (provider price)

View pricing

GPT Realtime 2.1 Mini

$0.02/month

$0.020 per minute (provider price)

View pricing

Phonic v1

$0.14/month

$0.140 per minute (provider price)

View pricing

Pros & Cons

Pros

  • Tests models on 82 real phone-call scenarios run three times each, not synthetic prompts
  • Scores reliability (pass-cubed), success rate, data accuracy, stalled calls, response time, and cost per minute
  • Every call is published with full transcript, tool calls, and scores
  • Head-to-head model comparison and a failure-analysis view of which scenarios break each model

Cons

  • Scenario set is narrow (appointment booking and Medicare intake), so results may not generalize to every use case
  • A benchmark, not a build tool — it helps you choose a model but does not create the voice agent
  • Per-minute pricing shown comes from each provider and can change independently
  • Some listed models have no published rate, leaving cost comparisons incomplete

Questions & Answers

Alternatives

Artificial Analysis, LMSYS Chatbot Arena, Vocode

Reviews & Ratings

—

0 reviews

5
0%
4
0%
3
0%
2
0%
1
0%

Sign in to rate and review Cekura Bench.

Sign in to review

No reviews yet. Be the first to review Cekura Bench!

Try Cekura Bench free