AI voice production is the work that happens between a script and a finished audio file: choosing a voice, picking a text-to-speech model, fixing what it mispronounces and checking every line before it ships. Onepin, from the speech-evaluation company Podonos, is an agent built to do that work automatically, and it now pitches itself at large-scale projects in 10 languages, from game dialogue and audiobooks to training courses and adverts.
Onepin does not make voices itself. It routes each line of a script to one of many commercial speech engines, such as ElevenLabs, OpenAI, Microsoft, Google, AWS and Cartesia, then scores the result with its own models and regenerates or repairs lines that fall short. Its tagline sums it up: “TTS makes voice. Onepin makes it production-ready.”
This article explains what Onepin is, how its pipeline works, what it costs, where its own claims disagree, who it is for and what UK teams should check before using it. For the state of the voice models it routes to, see our report on ElevenLabs Eleven v4.
Table of contents
- What Onepin Is: an AI Voice Production Agent, Not a TTS Model
- How Onepin’s AI Voice Production Pipeline Works
- Ten Languages and How Many Models? Checking Onepin’s Claims
- Pricing: What Large-Scale AI Voice Production Costs
- Why Validation Matters in AI Voice Production
- Who Onepin Is For: Large-Scale AI Voice Production
- How Onepin Compares With Going Direct
- What UK Teams Should Check
- Five Questions to Ask Before an AI Voice Production Trial
- AI Voice Production FAQ
- References and Further Reading
What Onepin Is: an AI Voice Production Agent, Not a TTS Model
Onepin’s own description is that it is “the agent layer for AI voice.” That distinction matters, because most AI voice production tools are either a model or an editor. Onepin sits above the models.
The pitch
Bring a script, and Onepin says it will “plan it, route each line to the best model, validate every line, fix what fails, and ship publish-ready audio.” It is explicit about what it is not: “Onepin is not a TTS model. It does not generate voice directly.” It also says it is not tied to any single speech provider.
Who built it
Onepin is built by the team behind Podonos, which the company describes as evaluation infrastructure for speech AI. Founder and chief executive Soohyun Bae recalls trying, during the pandemic, to turn favourite books into audio and finding that “voices mispronounced names, flattened the emotion, and fell apart the moment a sentence got complicated.” The founder’s conclusion: “the models were never the problem”; the missing piece was choosing, guiding and checking them.
The agent layer
Instead of wiring an AI voice production pipeline by hand, teams describe the job in plain language. Onepin says the agent “interviews you about the job, recommends voices per language, selects the model that fits your quality bar, latency budget, and cost ceiling,” and configures the workflow. Every step stays visible and editable, which is important for teams that need to explain how a piece of audio was made.
How Onepin's AI Voice Production Pipeline Works
Onepin’s homepage breaks every script into four steps. Each one targets a common failure in AI voice production.
Step 1: clean the text
A normaliser rewrites numbers, dates, prices, addresses and abbreviations into the words a narrator should say, so “$20” becomes “twenty dollars” and “St.” becomes “Saint” or “Street” depending on context. Onepin claims that “nearly half of the TTS errors are fixed by the normalizer.” That is a vendor figure, but anyone who has heard a voice model read “NW” as a word will recognise the problem, and why cleaning comes first in any AI voice production workflow.
Step 2: pronounce the hard words
Onepin says a four-million-word dictionary teaches the models how to say brand names, drug names and personal names, with examples such as AbbVie, Siobhán, semaglutide and Hermès. In AI voice production, getting one name wrong can make a whole line unusable, especially in healthcare and advertising.
Step 3: validate every line
Every line is scored for naturalness, word accuracy, clarity and pronunciation before it is delivered. Onepin’s homepage shows sample scores in the high 90s for the same sentence in English, Spanish, Japanese, Korean and German, and says the same bar applies in every language.
Step 4: fix what fails
Lines that miss the bar are regenerated with new settings, sent to a different model, or, where only one word is wrong, repaired in the existing audio. Onepin’s summary for AI assistants says the point is that “the bad lines in the run were identified and corrected before anything shipped.”
| Step | What Onepin does | Problem it targets |
|---|---|---|
| Clean | Normalises numbers, dates, prices, addresses, abbreviations | Misreads such as “NW” or “3042” |
| Pronounce | Applies a 4M-word dictionary and phonemes | Brand, drug and personal names |
| Route | Picks a model per line and language | No single model wins every language |
| Validate | Scores naturalness, accuracy, clarity, pronunciation | Bad takes slipping through |
| Fix | Regenerates, reroutes or repairs single words | Costly manual retakes |
The in-house models behind the scores
Onepin says five of its own models carry the pipeline: naturalness measurement calibrated per language, clarity measurement, pronunciation error detection, pronunciation error correction and voice similarity scoring. The interesting claim is about pronunciation. Checking a take by transcribing it back to text, Onepin argues, misses most mistakes, “because an ASR pass tends to return the word the system was trying to say.” Its detection model listens for the defect in the audio instead.
Ten Languages and How Many Models? Checking Onepin's Claims
For an AI voice production product built on measurement, Onepin’s own pages are surprisingly inconsistent. That does not mean the product is weak, but buyers should know which number applies to them.
Languages
Onepin’s FAQ and pricing table both say it supports 10 languages “today, with more coming”, and invite customers to email if they need another. Its summary for AI assistants says it works “in any language.” It does not publish the list of ten. The homepage demo shows English, Spanish, Japanese, Korean and German, so those five are a safe assumption; ask about the rest.
Models
The model count varies by page: “30+” on the homepage, FAQ and pricing page; 32 featured models from 19 providers on the models page; 24 in the summary line of its Python package on PyPI; and “100+” on its GitHub repository and AI summary. The likely explanation is that the larger figure counts every model and voice it can reach while the smaller ones count curated commercial engines, but Onepin does not say.
| Claim | Where it says one thing | Where it says another |
|---|---|---|
| Languages | 10 (FAQ, pricing) | “Any language” (AI summary) |
| Models | 30+ (homepage, FAQ); 32 featured (models page) | 24 (PyPI); 100+ (GitHub, AI summary) |
| Voice cloning | “On the way” (FAQ) | Automated cross-provider cloning described (AI summary) |
| Open-source models | “Not yet” (FAQ); “Coming soon” (pricing) | “Supports … natively” (AI summary) |
| GDPR | “Working toward GDPR compliance” (FAQ) | Not addressed elsewhere |
Number of TTS models Onepin claims, by page (bar length relative to 100)
Why it matters for large-scale projects
In small AI voice production jobs, the exact count does not matter. For a game studio voicing thousands of lines in several languages, it matters a great deal which engines are available in which language, whether cloning is live, and whether open-source models are an option. Get the answers in writing for your languages before signing up.
Pricing: What Large-Scale AI Voice Production Costs
Onepin charges in credits, at 100 credits to the dollar, and says one credit covers generation and validation together, with no separate speech-engine bill. Seats are free on every plan.
| Plan | Price (billed yearly) | Credits | Validated audio | Overage |
|---|---|---|---|---|
| Free | $0 | 1,000 one-time | About 15 min | None |
| Creator | $16 a month | 3,000 a month | About 100 min | $0.010 per credit |
| Studio | $80 a month | 15,000 a month | About 9 hours | $0.008 per credit |
| Scale | $240 a month | 50,000 a month | About 30 hours | $0.007 per credit |
| Enterprise | Custom | Custom, pooled | Custom | Custom |
Cost per validated hour
Using Onepin’s own estimates of validated audio, the effective price of AI voice production falls as plans grow. Creator works out at $16 for about 100 minutes, or $9.60 an hour. Studio is $80 for about 9 hours (540 minutes), or $8.89 an hour. Scale is $240 for about 30 hours, or $8.00 an hour.
Effective cost per hour of validated audio, US dollars (plan price ÷ Onepin’s validated-audio estimate)
Read the credit maths carefully
The Free plan’s 1,000 credits buy about 15 minutes, which is about 67 credits a minute. On the paid plans the same estimate is about 28 to 30 credits a minute (for example 15,000 credits ÷ 540 minutes is 27.8). Onepin does not explain the difference, and actual use will depend on which models a script is routed to, so run a real sample before budgeting a large AI voice production project.
What you own
Onepin says it assigns every output to the customer and only routes to engines that permit commercial use. Enterprise customers can negotiate an intellectual property indemnity. Downloads are not included on the Free plan, and bring-your-own-key, which lets you use your existing ElevenLabs or other provider accounts, starts on Studio.
Why Validation Matters in AI Voice Production
The case for Onepin rests on one idea: choosing a good model is not the same as knowing that every take is usable.
One bad line spoils a scene
In a 2,000-line game script or a 10-hour audiobook, even a small error rate means dozens of broken lines. Onepin claims that a model which tops a public leaderboard “may still fail 5-10% of outputs in production.” That is a vendor’s figure, but the principle is sound: in AI voice production at scale, checking every line is cheaper than finding the faults after release.
Transcription is a weak test
Many teams check generated audio by running speech recognition over it and comparing the text. Onepin’s argument, set out in a blog post titled “TTS accuracy: ASR is the wrong metric”, is that recognition models guess the intended word, so a misplaced stress or a mangled brand name still transcribes as correct. Its own detection model is built to hear those errors directly.
Consistency across languages
Large-scale AI voice production often means the same character or brand voice in several languages. Onepin says its naturalness scores are calibrated per language, so one quality bar can govern a multilingual run. That claim is hard to check from outside, but it is the right problem to solve.
Who Onepin Is For: Large-Scale AI Voice Production
Onepin names its target customers directly, and nearly all of them run AI voice production in volume.
Studios and publishers
Game studios voicing non-player characters across languages, audiobook publishers producing long-form narration, and localisation and dubbing teams with per-language quality checks are Onepin’s core market. Its Studio and Scale plans are named for them.
Learning, marketing and creators
E-learning platforms, advertising teams scaling voice creative across regions, YouTube and TikTok creators re-voicing content in other languages, and AI video makers adding voiceovers to footage from tools such as Google Veo or OpenAI Sora are also on the list.
Developers
Onepin has a Python SDK and command-line tool (pip install onepin). The first release appeared on PyPI on 28 May 2026 and the latest, version 0.16.1, on 17 September. Onepin says the package also installs an Agent Skill for coding assistants including Claude Code, Cursor, OpenAI Codex, Gemini CLI and GitHub Copilot, and that its agent works over MCP. Natural language processing turns a plain-English brief into a configured workflow, which is the agent’s main selling point for developers.
How Onepin Compares With Going Direct
Onepin’s compare page sets it against ElevenLabs, Speechify, Murf, OpenRouter, Together AI and fal.ai. The useful question for buyers is simpler: use one provider directly, use a model marketplace, or use an AI voice production layer like Onepin?
| Approach | Good for | Trade-off |
|---|---|---|
| One provider direct (for example ElevenLabs or Microsoft MAI-Voice) | Simple projects in well-supported languages | You do your own checking and retakes |
| Model marketplace (for example OpenRouter or fal.ai) | Developers who want many models behind one API | Access, not quality control |
| Production layer (Onepin) | Large-scale, multilingual work that must be right first time | Another vendor, credits pricing, 10 languages today |
Model prices underneath
The underlying engines vary widely in cost. Microsoft, for example, now charges $22 per million characters for MAI-Voice-2.1 and $15 for its Flash version. Onepin’s models page lists engines at prices ranging from 1,500 credits ($15) to more than 7,000 credits ($70) per million characters, which is why automatic routing on price as well as quality can cut AI voice production costs at scale. For an earlier look at Microsoft’s speech models, see our report on MAI-Transcribe-2.
What UK Teams Should Check
Onepin is a young AI voice production product, and some of the checks that matter most for UK buyers are not yet settled.
Data protection
Onepin says it is SOC 2 Type II compliant, runs regular penetration testing and does not train on customer data. But its FAQ says it is “working toward GDPR compliance”. Data processing agreements are listed only on the Enterprise plan. If scripts contain personal data, or you clone a real person’s voice, you need a UK GDPR-compliant agreement before you start, so treat this as a cybersecurity and compliance review, not just a procurement one.
Voice rights and consent
Voices are becoming a legal battleground. A Tokyo court recently held that human voices can be protected, as we reported in our story on voice rights in Japan, and California now requires disclosure of synthetic performers in some adverts. Use licensed voices, get written consent for any clone, and keep records of which engine produced which line.
Fit with your workflow
Large-scale AI voice production touches scripts, translation, review and delivery. Our workflow automation and AI strategy teams help organisations decide where an agent like Onepin should sit and which approvals must stay with people.
Five Questions to Ask Before an AI Voice Production Trial
Onepin’s free plan makes a test cheap, but a useful trial needs a plan. These are the questions we would put to Onepin, or to any AI voice production vendor, before committing a real project.
Which of the ten languages are live?
Ask for the exact list of languages and locales, and which engines serve each one. A language that only one engine supports gives the routing nothing to choose between.
Which engines will my lines use?
Ask whether you can see, per line, which model produced the audio. For brand voices and legal sign-off, a record of the engine behind every line is as important as the audio itself.
How are the scores set?
Onepin says you set the quality bar. Ask what the naturalness and clarity scores mean in practice, how they were calibrated for your languages, and whether you can export them with the files.
What happens to my scripts?
Scripts can contain unreleased product names, personal data or confidential training material. Ask where scripts are processed and stored, which subprocessors see them, and how long they are kept.
What does a retake cost?
In AI voice production, the real cost is often the second and third attempt. Ask whether regenerated or repaired lines use extra credits, and run a sample of your hardest lines to see.
AI Voice Production FAQ
What is Onepin?
Onepin is an AI voice production agent from Podonos. It turns scripts into finished audio by choosing a voice and a text-to-speech model for each line, scoring every line with its own quality models, and fixing or regenerating lines that fail.
What is AI voice production?
AI voice production is the process of turning a script into finished, publishable audio with AI: choosing voices and models, generating speech, checking it and fixing errors. Tools such as Onepin automate the checking and fixing, which is where most AI voice production time goes.
Which languages does Onepin support?
Onepin says it supports 10 languages today, with more coming. It has not published the full list; its homepage demo shows English, Spanish, Japanese, Korean and German.
Does Onepin make its own voices?
No. Onepin routes to commercial speech engines from providers such as ElevenLabs, OpenAI, Microsoft, Google, AWS and Cartesia, and adds cleaning, pronunciation, validation and correction on top.
How much does Onepin cost?
There is a free plan with 1,000 one-time credits. Paid plans billed yearly are $16, $80 and $240 a month, with 3,000, 15,000 and 50,000 monthly credits. Enterprise pricing is custom.
Can I use my own ElevenLabs account?
Yes, on the Studio plan and above, through bring-your-own-key. Onepin then adds routing and validation on top of your existing provider account.
Who owns the audio?
Onepin says it assigns every output to the customer and only uses engines that allow commercial use. Enterprise customers can negotiate an IP indemnity.
References and Further Reading
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.