Sources
See it in action
Browse the models and styles behind stories like this one — free account, instant gallery.
Explore the catalogDiscuss this with
Pick a companion and get their take on this story

Sofia follows the money, policy, and platforms shaping what creators can make.
Browse the models and styles behind stories like this one — free account, instant gallery.
Explore the catalogPick a companion and get their take on this story
Hugging Face has launched the Open TTS Leaderboard, a scalable public benchmark that ranks text-to-speech and voice cloning models across multiple languages and evaluation dimensions — giving AI creators their first apples-to-apples comparison of the systems increasingly powering AI companions, narration, and character voice work.
Until now, choosing a voice model for an AI character or narration workflow meant trusting a provider's own demos — a situation roughly analogous to buying paint based on the manufacturer's swatch. The Open TTS Leaderboard, detailed on the Hugging Face blog, changes that by publishing standardized scores across three axes that actually matter in production: naturalness (does it sound human?), intelligibility (can listeners accurately transcribe it?), and speaker similarity (does voice cloning actually match the target speaker?).
The multilingual scope is the detail that separates this from earlier efforts. Most TTS shootouts have been English-centric, which tells creators almost nothing about how a model handles French character dialogue, Japanese narration, or Spanish companion voices. Hugging Face's framework tests across languages systematically, which matters as AI-companion platforms and character generators push into non-English markets.

The Open TTS Leaderboard ranks TTS and voice cloning models on naturalness, intelligibility, and multilingual performance.
Image: Hugging Face Blog
The leaderboard pairs automated metrics with human preference evaluations. That dual approach is a direct response to a known failure mode: models can be tuned to score well on signal-level metrics like character error rate while still sounding robotic or unnatural to a human ear. By anchoring rankings to both, Hugging Face makes it harder for a model to game its way to the top without genuinely sounding good.
For creators building AI companions or generating character voices, this distinction is practical. A model that tops an automated-only chart might still produce the flat, slightly-off cadence that breaks immersion in a roleplay scene or an audio narration. Human preference scores surface that gap.
"Scalable evaluation is key to driving progress in TTS — we need benchmarks that work across languages and capture what humans actually care about."
— Hugging Face Blog
The leaderboard accepts external model submissions, which is where the power dynamic gets interesting. When Stability AI faced community pressure over model quality claims in 2023, the absence of neutral third-party benchmarks made independent verification nearly impossible. An open, submittable leaderboard removes that bottleneck: a small lab or individual researcher who fine-tunes a multilingual voice model on a niche language can now put it on equal footing with a commercial provider's flagship system.
For creators who fine-tune their own TTS models — or who rely on open-weight voice systems for character generation — that means the leaderboard becomes a living resource, not a one-time snapshot. Rankings will shift as new submissions arrive, so checking it before committing to a voice backend for a long-running project is worth building into the workflow.
Creators already experimenting with AI character voices through Charmloop's generation tools or exploring voice-adjacent model options in the model catalog will find the leaderboard a useful external reference point when evaluating which TTS system to pair with their characters.
The leaderboard is live now on Hugging Face. The immediate question for the field is which major commercial TTS providers — ElevenLabs, Cartesia, PlayHT, and others — choose to submit their models for independent ranking, and which stay on the sidelines. That decision, more than any score, will signal how confident each provider is in their own product.