Evidence snapshot · Evaluated 2026-07-24

18 Voice Candidates, Tested Before Recommendation

We ran 126 automatic synthesis checks across 18 Kokoro-82M voices, then used a fixed listening rubric to decide which voices deserved a featured position. Eight English voices passed. The tested Chinese voice set did not.

Results at a Glance

126 / 126
automatic audio checks passed
8 / 18
candidates passed the featured gate
0 / 8
Chinese candidates passed the Chinese launch gate
0 / 18
long-form samples with a perceived voice switch

Download the public results as JSON.

How to Read the Scores

Each score is from 1 to 5 in this order: naturalness, pronunciation accuracy, long-sentence stability, long-form consistency, audio integrity, and Chinese-English code-switching. English code-switch scores are diagnostic only because these voices are recommended as English voices. Mixed-language speech is a release gate for Chinese candidates, so cross-language averages would be misleading and are not published.

The listening review was completed by one internal reviewer using headphones in an ordinary room. This is a product quality gate, not an independent laboratory study or a statistically representative preference test.

Candidate Results

VoiceLocaleSix scoresGate resultRecommended scenesLive short preview
Aoede
af_aoede
en-US 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
course, podcast, article
Bella
af_bella
en-US 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
article, course, podcast
Kore
af_kore
en-US 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
article, course, podcast
Nicole
af_nicole
en-US 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
article, podcast
Fenrir
am_fenrir
en-US 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
podcast, article, course
Michael
am_michael
en-US 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
podcast, video
Emma
bf_emma
en-GB 5 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
course, article, podcast
Puck
am_puck
en-US 5 / 5 / 5 / 5 / 5 / 1 Did not pass
no use-case score reached 4/5
None
Heart
af_heart
en-US 4 / 5 / 5 / 5 / 5 / 1 Featured
Passed every applicable featured gate
article
Yunjian
zm_yunjian
zh-CN 4 / 5 / 5 / 5 / 5 / 1 Did not pass
blocking listening issue, Chinese-English code-switch score below 3/5
article, course
Xiaobei
zf_xiaobei
zh-CN 4 / 4 / 4 / 4 / 4 / 1 Did not pass
blocking listening issue, Chinese-English code-switch score below 3/5
podcast
Xiaoni
zf_xiaoni
zh-CN 4 / 4 / 4 / 4 / 4 / 1 Did not pass
blocking listening issue, Chinese-English code-switch score below 3/5, no use-case score reached 4/5
None
Sarah
af_sarah
en-US 3 / 5 / 5 / 5 / 5 / 1 Did not pass
naturalness score below 4/5
article, podcast
Xiaoxiao
zf_xiaoxiao
zh-CN 3 / 5 / 5 / 5 / 5 / 1 Did not pass
blocking listening issue, naturalness score below 4/5, Chinese-English code-switch score below 3/5
article, course, podcast
Yunyang
zm_yunyang
zh-CN 3 / 4 / 4 / 5 / 5 / 1 Did not pass
blocking listening issue, naturalness score below 4/5, Chinese-English code-switch score below 3/5
article
Yunxia
zm_yunxia
zh-CN 3 / 4 / 4 / 5 / 5 / 1 Did not pass
blocking listening issue, naturalness score below 4/5, Chinese-English code-switch score below 3/5, no use-case score reached 4/5
None
Xiaoyi
zf_xiaoyi
zh-CN 3 / 4 / 4 / 5 / 4 / 1 Did not pass
blocking listening issue, naturalness score below 4/5, Chinese-English code-switch score below 3/5
course, podcast, video, article
Yunxi
zm_yunxi
zh-CN 3 / 4 / 3 / 5 / 5 / 1 Did not pass
blocking listening issue, naturalness score below 4/5, Chinese-English code-switch score below 3/5, no use-case score reached 4/5
None

Live previews use the current production preview sentence and are not the original benchmark recordings.

Why the Chinese Launch Was Paused

All eight tested Chinese voices passed automatic audio-integrity checks, but none met the featured-quality gate. Every candidate failed the Chinese-English mixed-speech requirement, and the listening reports also recorded blocking pronunciation or abbreviation issues. Because the launch condition required at least four featured-quality Chinese voices, the Chinese-market validation phase did not start.

This is a provider-quality result, not a demand result. It does not show that Chinese creators do not want document-to-audio software; it shows that the tested voice set was not strong enough to run a fair demand test.

Limitations and Reproducibility

  • One reviewer means the scores should be treated as a release decision, not universal preference.
  • Automatic checks verified that audio was generated and structurally valid; they did not prove naturalness.
  • The candidate set was deliberately narrow: ten English voices and all eight available Chinese voices.
  • Provider, model version, evaluation date, criteria, and public result data are disclosed so later runs can be compared honestly.

Read the full evaluation and comparison-page methodology.