Suno Speech is the AI music company’s first product that is not mainly about songs. Launched in public beta on 1 October 2026, Suno Speech generates spoken-word audio and an original backing score together, in one track, from a typed idea or a full script. Suno calls it “the first audio model that generates voice and music together as one cohesive track.”
The beta is open to every Suno user inside the normal app on web, iOS and Android, and it lands three weeks after Suno’s v6 music models. It also puts Suno up against voice specialists such as ElevenLabs, whose own new voice models arrived days earlier. For background on the company, see our guide to what Suno AI is.
This article covers what Suno Speech does, how to use it, what Suno has and has not said, why the company is branching out, and what it means for businesses thinking about synthetic voice in marketing, training or customer content.
Table of contents
- What Suno Speech Is
- How to Use Suno Speech
- What Suno Speech Can Make
- Beta Really Does Mean Beta
- Getting Better Results From Suno Speech
- What Suno Has Not Said
- Why Suno Is Moving Beyond Songs
- Suno Speech vs the Voice Specialists
- The Risks Worth Weighing
- What Suno Speech Means for Businesses
- Suno Speech FAQ
- References and Further Reading
What Suno Speech Is
The short version: you type words or an idea, describe a voice and a musical mood, and Suno Speech returns a spoken performance with a soundtrack underneath. The two layers are generated together rather than recorded separately and mixed afterwards.
Suno’s own description
Suno’s announcement, written by chief product officer Jack Brody and published at 20:35 UTC on 1 October, says Suno Speech “lets you create spoken audio set to original background music.” It adds: “type an idea, a poem or something you’ve written, then describe the voice and musical style you have in mind.”
The release note posted the same day is shorter and more playful: “Meet the first model to create a speech and its soundtrack together, in one take. Make bedtime stories over soft piano, hype speeches over stadium drums, ASMR grocery lists and more.” It is tagged for Android, iOS, web and the Create screen.
Creative entertainment, not just tools
Brody frames Suno Speech as part of what Suno calls creative entertainment: people getting fulfilment from “the simple act of making something.” He writes that “music will always be at the heart of Suno,” but that its vision “has always extended to other forms of human expression.”
How Suno’s team used it
The post describes the team’s own experiments while building Suno Speech. “We turned friends’ texts into wildly overproduced dramatic readings. We gave ordinary voice notes unnecessarily epic scores. We made meditations, poems, pep talks, and bedtime stories for our kids.” That list is a good guide to the use cases Suno expects.
How to Use Suno Speech
Suno has not published a help article for Suno Speech yet, so the clearest walk-through so far comes from The Verge’s Jess Weatherbed, writing on 2 October.
Finding the feature
Open the Create tab and choose the Speech option. Suno Speech sits alongside song creation rather than in a separate app, which matches the company’s line that it is “built right into Suno.”
Simple mode
Simple mode works like Suno’s song prompts. You describe what you want in a prompt box, such as “a pirate captain rallying his crew”, and Suno Speech writes and performs the words itself, with a matching score.
Advanced mode
Advanced mode is for people who already know the exact words. You paste a custom script, then adjust the gender of the AI voice, the speech style, and how much variety each generation has. According to The Verge, a Suno Speech track can run to around eight minutes.
Turning the music off
The soundtrack is optional. A toggle switches off the background music if you only want clean speech, which makes Suno Speech usable as a straight text-to-speech tool too.
| Setting | Simple mode | Advanced mode |
|---|---|---|
| Input | A short description of the idea | Your own script |
| Who writes the words | Suno | You |
| Voice controls | Described in the prompt | Voice gender, speech style, variety |
| Background music | On by default, can be switched off | On by default, can be switched off |
| Maximum length | Around eight minutes | Around eight minutes |
What Suno Speech Can Make
Suno’s examples point at personal, playful content first, with obvious spill-over into podcasts, video and marketing.
Personal and family audio
Bedtime stories over soft piano are Suno’s headline example, and Brody notes that people already make songs “for birthdays, weddings, inside jokes, faith and worship, their kids, their friends.” Suno Speech gives that same audience a spoken format: a recorded toast, a message for a grandparent, or a story with a child’s name in it.
Dramatic and comic readings
The team’s favourite trick was turning ordinary text into high drama, and that is likely to be how Suno Speech spreads on social media. A grocery list read as an ASMR whisper or a group-chat argument read like a film trailer is exactly the kind of clip people share.
Speeches, pep talks and meditations
Hype speeches over stadium drums and calm meditations are the two ends of the mood range Suno advertises. Because voice and score come from one model, the pacing of the music can follow the delivery instead of running underneath it at a fixed tempo.
| Example from Suno | Voice | Music |
|---|---|---|
| Bedtime story | Soft, slow narration | Soft piano |
| Hype speech | Loud, rallying delivery | Stadium drums |
| ASMR grocery list | Whispered | Ambient |
| Friend’s text as drama | Theatrical reading | “Wildly overproduced” score |
| Meditation or poem | Calm, measured | Gentle backing |
Beta Really Does Mean Beta
Suno is unusually frank about the rough edges, and the warnings are worth reading before anyone puts Suno Speech near a client.
Accents that wander
“Occasionally, British accents can wander off to Australia and back,” Brody writes. For UK businesses that is the most practical warning in the post: an accent that drifts mid-sentence is fine for a joke clip and not fine for a brand voice.
Very dramatic pauses
“Dramatic pauses may be very dramatic,” the post continues. Timing is one of the hardest parts of synthetic speech, and a model that also has to fit music around the words has more ways to get it wrong.
A month of testing first
Suno says it spent “the past month testing Speech with a small group of users” before opening the beta to everyone, and that it will keep improving Suno Speech “as we learn what people love, what doesn’t quite work yet, and what our community wants to do next.”
Getting Better Results From Suno Speech
Suno has not published a prompting guide for Suno Speech yet, but the controls The Verge describes, and the warnings in Suno’s own post, suggest a sensible way to work with the beta.
Write for the ear
Scripts written for reading rarely sound natural aloud. Short sentences, plain words and one idea per line give Suno Speech less room to stumble, and they make any odd emphasis easier to spot and fix. Read the script aloud yourself once before generating it.
Name the accent and check it
Because Suno itself warns that British accents can drift, say exactly what you want in the voice description, such as a calm southern English accent, and listen to the whole track rather than the first few seconds. If the accent wanders, regenerate rather than editing around it.
Describe the music in plain terms
The soundtrack follows the musical style you describe. Genre, instruments and energy level, such as soft piano at a slow tempo or stadium drums that build, give Suno Speech something concrete to work with. If the music competes with the words, switch it off and add your own backing later.
Use Advanced mode for exact wording
Simple mode writes the words for you, which is fun but risky for anything factual. For product names, prices or instructions, paste your own script in Advanced mode so Suno Speech cannot invent details. Use the variety setting to generate a few takes and keep the best one.
Keep long pieces in sections
With a ceiling of around eight minutes, longer material has to be split anyway. Generating section by section also means a single bad pause or accent slip costs you one short take, not the whole recording.
| Problem | What to try |
|---|---|
| Accent drifts mid-track | Name the accent precisely, then regenerate |
| A pause runs too long | Shorten the sentence before it, then regenerate |
| Music drowns out the voice | Ask for softer, sparser music, or switch it off |
| Simple mode invents details | Use Advanced mode with your own script |
| Script is longer than eight minutes | Split it into sections and generate each one |
What Suno Has Not Said
The announcement, the release note and The Verge’s report leave several practical questions open. We checked each of them against the published material on 2 October.
Price and credits
None of the three sources says how many credits a Suno Speech generation costs, or whether free and paid users get different limits. Suno’s song generation runs on credits, with advanced tools such as Voices and Suno Studio reserved for Pro and Premier plans, so a credit cost is likely, but it is not stated.
Your own voice
Suno already offers Voices, a Pro and Premier beta that lets you record your own singing and use it in songs, with what Suno’s own guide calls “a verification step to confirm the voice is yours.” Nothing published so far says whether Voices can be used with Suno Speech, or whether the speaking voices are drawn only from Suno’s own range.
Languages, rights and labelling
Suno has not listed which languages Suno Speech supports, what commercial rights paid users get over speech tracks, or whether the audio carries a watermark or other marker identifying it as AI-generated. All three matter for business use, and all three are worth checking in Suno’s terms before publishing anything.
| Question | Answer so far | Source |
|---|---|---|
| Who can use it? | Everyone, in beta, on web and mobile | Suno blog, release note |
| Maximum length | Around eight minutes | The Verge |
| Music optional? | Yes, by toggle | The Verge |
| Credit cost | Not stated | – |
| Works with your own voice? | Not stated | – |
| Languages | Not stated | – |
| AI label or watermark | Not stated | – |
Why Suno Is Moving Beyond Songs
Suno Speech is a product launch, but it is also a business decision, and the timing tells part of the story.
Diversifying away from the lawsuits
The Verge’s reading is blunt: Suno is “likely” branching out “in an attempt to diversify the platform, given its music generator has attracted so many lawsuits.” Warner Music Group settled with Suno and became a partner in November 2025, but as of September, Sony Music and Universal Music were still pursuing their case against the company in a federal court in Massachusetts.
Licensed music models arrive first
Three weeks before Suno Speech, Suno launched v6, a family of music models it says were “developed with” Warner Music Group, BMG and Believe. We looked at how carefully Suno worded that launch in our report on Suno v6 licensed music. The flagship v6 and the experimental v6-wild are for Pro and Premier subscribers, and v6-mini is available to everyone.
Money to spend
In June 2026 Suno said it had raised more than $400 million in a Series D round at a $5.4 billion post-money valuation, led by Bond Capital. Chief executive Mikey Shulman wrote that viral trends had taken Suno to number one in the App Store’s music category “in dozens of countries.” That cash pays for new models, and Suno Speech is the first to point the company at a market beyond music.
Days between Suno’s major 2026 launches (bar length relative to 140 days)
The gaps are counted from the dates in Suno’s own release notes. 26 March to 13 August is 140 days, 13 August to 9 September is 27 days, and 9 September to 1 October is 22 days. Bars are scaled to the longest gap, so 27 ÷ 140 = 19.3% and 22 ÷ 140 = 15.7%. Suno is shipping big launches faster now than it was in the spring.
Suno Speech vs the Voice Specialists
AI speech is not new, and Suno Speech enters a crowded field. Its pitch is the combination, not the voice alone.
A decade of synthetic speech
The Verge points out that DeepMind has been working on deep-learning speech synthesis for about ten years, that Adobe offers a text-to-speech tool, and that ElevenLabs has become one of the best-known names since launching in 2023. All of these turn text into a voice using methods from natural language processing and audio modelling.
ElevenLabs moves the same week
ElevenLabs launched Eleven v4 and v4 Turbo on 28 September, three days before Suno Speech, and they went straight to the top of a public text-to-speech leaderboard. We covered that release in our report on ElevenLabs’ new voice models. ElevenLabs also sells a separate music model, so a creator can already get voice and music from one company, just not from one generation.
Where Suno Speech is different
The difference is in the workflow. With most tools you generate a voiceover, find or generate music, then line the two up in an editor. Suno Speech does that in one step, which suits short personal and social clips. For long, precise work such as an audiobook or an e-learning course, dedicated voice tools still offer more control today.
| Need | Suno Speech | Dedicated voice tool |
|---|---|---|
| Voice and music in one pass | Yes, by design | Usually two steps |
| Clean speech only | Yes, music toggle off | Yes |
| Fine control of timing and pronunciation | Limited in beta | Usually stronger |
| Long-form audio | Up to around eight minutes | Often much longer |
| Maturity | Public beta | Generally available |
The Risks Worth Weighing
Synthetic voice raises questions that synthetic music mostly does not, because a voice can stand in for a person.
Impersonation and consent
A voice that sounds like a real person can mislead listeners, whether the target is a celebrity, a chief executive or a family member. Suno’s Voices feature already uses a verification step for singing voices. Whether Suno Speech has equivalent safeguards against imitating a real speaker has not been described, and it is the first thing a careful buyer should ask.
Disclosure
Listeners increasingly expect to be told when a voice is synthetic, and the EU AI Act includes transparency rules for AI-generated audio. Even where the law is unsettled, clear labelling protects trust. A short line in a video description or podcast note costs nothing.
Rights in the score
Because Suno Speech generates the music as well as the voice, the music questions that have followed Suno for two years apply here too. The v6 deals with labels change the picture for Suno’s newest models, but businesses should still read the terms for commercial use before putting a Suno Speech track in an advert.
What Suno Speech Means for Businesses
For most organisations, Suno Speech is a quick way to prototype audio, not yet a production voice platform.
Good early uses
Internal training snippets, social video intros, event announcements, podcast stings and mock-ups for agency pitches are all sensible tests. Each is short, low-risk and easy to redo if an accent wanders or a pause runs long. Suno Speech can turn a script into a scored draft in minutes, which is useful even if the final version is recorded by a person.
Where to hold back
Customer-facing material that carries your brand voice, anything that could be mistaken for a real person speaking, and regulated content such as financial promotions deserve more caution while Suno Speech is in beta. Keep a human review step, and keep the script and settings so the audio can be reproduced or corrected.
Fit it into a wider AI plan
Generative audio is one more tool in a growing stack, and it works best with clear rules on approval, labelling and data. Our AI strategy team helps businesses set those rules, and our AI models and tools hub tracks the releases worth testing.
Suno Speech FAQ
What is Suno Speech?
Suno Speech is a beta feature in Suno that generates spoken-word audio and original background music together in one track, from a short idea or a full script.
When did Suno Speech launch?
Suno announced Suno Speech on 1 October 2026 and opened the beta to all users on web, iOS and Android after a month of testing with a small group.
How long can a Suno Speech track be?
According to The Verge, up to around eight minutes.
Can I use Suno Speech without music?
Yes. A toggle turns off the background music for clean speech.
How much does Suno Speech cost?
Suno has not said how many credits a generation uses or whether limits differ by plan.
Is Suno Speech good enough for professional voiceovers?
It is a beta, and Suno itself warns that accents can drift and pauses can run long. It is well suited to drafts and casual content. For polished brand or long-form work, test it carefully against dedicated voice tools first.
References and Further Reading
Introducing Speech (beta), by Jack Brody (Suno blog)
Suno release notes: Introducing Speech (beta), 1 October 2026
AI music maker Suno now generates spoken words (The Verge)
Suno Launches Speech Beta, Pairing Voice and Music in One Model (Unite.AI)
Introducing v6, by Mikey Shulman (Suno blog)
The Next Chapter for Suno: Series D announcement (Suno blog)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.