Suno Vocal Prompts: Separating Voice Identity From Vocal Performance

Three complaints dominate the v6 vocal discussion in its first week: the singer rushes the lyric, the delivery sounds spoken rather than sung, and on two-voice songs the singers swap parts mid-line. They look like one problem and they are not. Two of them respond to how you write the lyric, one responds to structure, and none of them respond to the thing most people try first — piling more adjectives into the style box.

Filed 2026-09-11 Read 9 min Method How we work
In short
  • Voice identity (who is singing) and vocal performance (how they sing it) are controlled by different parts of your input. Style tags set identity. Lyric structure sets performance. Most failed vocal prompts are performance problems being addressed with identity tools.
  • Rushed delivery is usually a syllable-density problem, not a prompt problem. A line with too many syllables for its bar leaves the model no room to sustain anything, so it compresses. Shorter lines are the most reliable lever.
  • Users report that spelling a vowel out to signal a held note helps, but this is a community technique from v6's first days, not a verified behaviour. Treat it as worth testing on your own material rather than as a documented feature.
  • Duet failures split into three different problems — alternating verses, call-and-response, and simultaneous harmony — and they need different lyric structures. Reports on whether v6 handles them are genuinely contradictory right now.
  • The Variety slider rewrites your style prompt unless you set it to 0. If your vocal tags keep getting ignored, check that first — it is a settings problem wearing a prompting problem's clothes.
  • Nothing about a good vocal changes distributor screening. Raw AI tracks were rejected 50 of 50 by DistroKid, 47 of 50 by TuneCore and 42 of 50 by CD Baby in our benchmark, and a better performance is not a cleaner file.
Three short spacious lyric lines in pink beside one long crowded grey line, representing how syllable density decides whether a vocal sings or sounds spoken
Shorter lines create the room a held note needs. No style tag can do that.

Good suno vocal prompts start by deciding which of two separate things you are trying to change. Almost every failed vocal in Suno v6's first week comes from getting this wrong, and the confusion is completely reasonable, because the interface does not draw the line anywhere.

Voice identity is who is singing: a breathy female alto, a gravelly baritone, a choir. That lives in your style tags.

Vocal performance is how they sing it: where they breathe, which words they hold, whether they rush the third line. That lives almost entirely in the shape of your lyric, and only marginally in your style tags.

When someone writes "emotional, expressive vocals, powerful delivery" into the style box and gets back a singer who rattles through the verse like a train announcement, they have described an identity and expected a performance. The model did what was asked. It just was not asked the right thing.

This guide covers the three vocal failures dominating v6's launch discussion — rushed and spoken delivery, notes that will not sustain, and duets where the singers swap parts mid-line — and separates what prompting genuinely reaches from what it does not.

A note on how new this evidence is

Suno released v6 on 9 September 2026. This page was written on 11 September. Everything here that comes from the community is therefore two days old, and some of it directly contradicts other parts of it: one producer's detailed write-up reports duet conversion working cleanly in Simple mode, while another thread from the same window reports duets failing outright.

Voice identity and vocal performance are set in different places — Describing an identity and expecting a performance is why 'emotional, expressive vocals' returns a singer who rushes the verse.
Style tags cannot create room in a lyric that has none.

We are not going to paper over that. Where a technique is documented by Suno, we say so and link it. Where it is a community finding from v6's first days, we label it as one and tell you how to test it yourself. Our full take on the release is in the Suno v6 review.

Rushed and spoken delivery is a syllable problem

This is the most common complaint and the one with the least intuitive fix.

A vocal line has a fixed amount of musical room. If your lyric puts twelve syllables where the phrase has space for seven, something has to give, and what gives is duration — each syllable gets shortened until they all fit. Shortened syllables, strung together, are what speech sounds like. The model is not refusing to sing. It is singing a line that has no room to be sung.

So the lever is the lyric, not the adjective:

The vowel-spelling technique. Several users in v6's first week report that writing a held word phonetically — "staaay" rather than "stay" — produces a sustained note where explicit delivery instructions did nothing. One detailed account describes this working after a long run of failed attempts using conventional instructions.

We are flagging this carefully. It is a plausible mechanism, it is reported by more than one person, and it costs nothing to try. It is also two days old, it is not in Suno's documentation, and one user's run of attempts is not a benchmark. Test it against your own material before you build a workflow on it.

Duets are three different requests

"Make it a duet" is the single most overloaded instruction in AI music, because it can mean three structurally different things, and the model has to guess which one you meant.

'Make it a duet' can mean three different things — The model has to guess which one you meant. Asking for the specific behaviour is what stops singers swapping mid-line.
Reports on how reliably v6 handles each are contradictory in its first week.

Decide explicitly, then write the lyric to match:

Alternating verses. Two voices taking turns across whole sections. This is a structural request and it needs structural markup — label every block, and never split a sentence across a label boundary. The model infers voice assignment from structure, so a continuous unlabelled block gives it nothing to respect.

Call and response. A short phrase answered by a second voice. This needs short paired lines, written as pairs, with the answer clearly shorter than the call. Long call-and-response lines collapse back into alternating verses.

Simultaneous harmony. Both voices on the same words at the same time. This is a texture instruction rather than a structural one, and it belongs in your style tags rather than in the lyric layout.

Reports on how well v6 handles any of these are, right now, genuinely contradictory. One producer's write-up describes singers finally harmonising without swapping mid-sentence. Another thread from the same two days reports duets still failing. Both are honest reports from people with access to different models in a staggered rollout, which is a good reminder that launch-week consensus does not exist yet.

What is stable advice regardless of which report is right: label every part, keep labelled blocks to whole lines, and ask for one of the three behaviours rather than for "a duet".

Intros and endings are a structure problem

Two recurring complaints belong here rather than with delivery: unwanted humming or vocalising before the first line, and songs that stop mid-lyric.

Both are section-boundary problems. An unrequested intro vocal usually means the model has been given a song shape with a gap at the front and has filled it. Writing an explicit opening — an instrumental intro instruction, or a first line that starts immediately — gives it something to put there instead. Endings that cut off mid-line usually mean the final section had more lyric than remaining length; shortening the last line, or giving the song an explicit outro to land in, resolves most cases.

Where prompting does not resolve it, the edit tools in v6 will. This is the generation where you can change one section without rebuilding the track, which makes a bad ending a thirty-second fix rather than a re-roll.

When your vocal tags are being ignored entirely

If your style tags seem to have no effect at all — not a weak effect, no effect — stop rewriting the prompt and check your sliders.

Suno's own v6 FAQ states that the Variety slider works by "adjusting and updating your style prompts", and that to retain full control of your style tags you should reduce Variety to 0. That is a documented behaviour, not a theory, and it means a non-zero Variety setting can be rewriting the vocal tags you are carefully tuning.

This trips up enough people that it has its own page: Suno style influence, Variety and Weirdness explained covers each control separately and what the documentation actually says about it.

A worked rewrite

The abstract advice becomes obvious once you see the same idea written twice.

The version that comes back spoken:

I never thought that I would see you standing there in my doorway again on a cold night like tonight

Twenty-four syllables in a line built for roughly half that. There is no possible performance of this that is not rushed, and no style tag repairs it.

The version that sings:

I never thought I'd see you here

Standing in my door

On a night this cold

Three short lines, each ending on a word with space after it — "here", "door", "cold". Every one of those is now available to be held. Nothing was added to the style prompt; the room was created in the lyric.

The same lines with sustain signalled, using the community technique described above:

I never thought I'd see you heeere

Standing in my door

On a night this coooold

Generate all three against an identical style prompt with Variety at 0, and you have a controlled test of both the structural change and the vowel technique on your own material — which is worth more than anyone's assurance, including ours.

A quick diagnostic

Symptom Most likely cause Where to fix it
Lyrics rushed, sounds spoken Too many syllables per line Shorten the lyric
Notes never sustain Key word sits mid-line Move it to the end of a phrase
Singers swap mid-sentence No structural boundary Label parts, keep blocks whole
Voice is the wrong character Identity, not performance Style tags
Style tags ignored entirely Variety slider rewriting them Set Variety to 0
Vocal buried under the band Arrangement density Thin the arrangement
Vocal dull on some rolls only The render, not the prompt Re-roll, or stems

Where prompting stops

Being honest about the ceiling saves more time than any technique on this page.

Prompting reaches lyric structure, section boundaries, voice identity, and arrangement density. It does not reach the rendered audio. If your vocal is dull, thin, or sits behind the instruments because of how the model produced the file rather than because of what you asked for, no rewording fixes it — that is a mix problem. Your options there are stems and a mastering pass, which our Suno stems guide covers, or regenerating and hoping for a better roll.

Telling the two apart is simple in practice. If the same prompt produces the problem consistently across several generations, it is your prompt. If it appears on some rolls and not others with identical input, it is the render, and you are re-rolling rather than rewriting.

The step after the vocal is right

A vocal you are finally happy with is still an AI-generated file, and that is a completely separate gate from whether it sounds good.

Distributor ingestion screening does not grade performances. It scans for the statistical fingerprint that generation leaves in the audio, plus the embedded provenance data Suno now adds — and v6 was built with Warner Music Group, BMG and Believe, so that trail is getting stronger rather than weaker. In our benchmark on raw, unprocessed AI tracks, DistroKid rejected 50 of 50, TuneCore 47 of 50 and CD Baby 42 of 50, against approximate confidence thresholds of 0.78, 0.82 and 0.85.

Undetectr works on Suno v6 output and is the tool in our benchmark that addresses that layer. It processes the audio signal across six layers — spectral artifact removal, temporal pattern normalisation, dynamic range processing, metadata sanitisation, and the removal of embedded SynthID and C2PA watermarks — which is why it carries across a model generation rather than breaking on each new release. Those are the artifacts DistroKid, Spotify, Apple Music and Amazon Music actually scan for.

The order that works: get the vocal right, then clean the file, then distribute. Our guide to uploading Suno tracks to Spotify covers the distribution half.

How we checked

The Variety slider behaviour is quoted from Suno's own v6 FAQ, fetched on 11 September 2026, which states that Variety adjusts and updates style prompts and that reducing it to 0 retains full control of style tags. The Weirdness and Style Influence definitions come from Suno's Creative Sliders documentation, fetched the same day.

Community reports come from a structured pull of Reddit and YouTube discussion across the window 12 August to 11 September 2026 — 25 Reddit threads carrying 782 upvotes and 1,351 comments, plus six YouTube videos with transcripts. Because v6 launched on 9 September, direct v6 experience in that pool covers two days.

Three caveats. Contradictory reports in the same window are reported as contradictory rather than resolved in favour of whichever is tidier. Techniques we label as community findings have not been independently tested by us on a controlled corpus. And the distributor rejection figures describe a pre-v6 corpus; that re-run is in progress and will be published when it completes.

Frequently asked

Questions readers ask.

Most often because the lyric line has more syllables than the phrase has room for. When the model has to fit twelve syllables into a bar built for seven, the only way through is to compress each one, and compressed syllables read as speech rather than song. The fix is structural rather than descriptive: cut the line down, break it across two lines, or leave deliberate space at the end. Adding words like 'sung expressively' to your style prompt does not create room that the lyric does not have.

Give the note somewhere to live by shortening the line around it, and put the word you want sustained at the end of a phrase where nothing follows it. Users in v6's first week also report that spelling a vowel out — writing a held 'stay' as 'staaay' — signals duration, and several say it worked where explicit delivery instructions did not. We are flagging that as a community technique rather than documented behaviour, because it has had days rather than months of testing. It costs nothing to A/B on your own track.

Because nothing in the lyric tells it where one singer stops. Suno infers voice assignment from structure, so if your lyric is a continuous block the model has no boundary to respect. Label every part explicitly and keep each labelled block to whole lines rather than splitting a sentence across two singers. Reports on how reliably v6 respects those labels are contradictory right now — some users report harmonies finally holding without mid-sentence swaps, others report duets still failing outright.

They are three different requests and conflating them is why duet prompts fail. Alternating verses means two voices taking turns, which needs labelled blocks. Call-and-response means a short phrase answered by a second voice, which needs short paired lines. Simultaneous harmony means both voices on the same words at once, which is a texture instruction rather than a structural one. Ask for the one you actually want; asking for 'a duet' leaves the model to pick.

Partly. If the vocal is buried because the arrangement is dense, thinning the instrumentation in your prompt genuinely helps, and our Suno arrangement prompts guide covers that. If the vocal is dull because of how the model rendered the audio itself, prompting cannot reach it — that is a mix problem and it belongs in a mastering pass or in stems. Knowing which one you have saves a lot of wasted generations.

No. Distributor ingestion screening looks for the statistical fingerprint and embedded watermarks that generation leaves in the audio, not for musical quality. Our raw-corpus figures are 50 of 50 rejected by DistroKid, 47 of 50 by TuneCore and 42 of 50 by CD Baby. Undetectr works on Suno v6 output and removes those artifacts — spectral signatures, timing grids, and embedded SynthID and C2PA provenance data — so the file clears screening on Spotify, Apple Music and the rest.

The verdict, in one sentence: Undetectr.

A vocal you are finally happy with is still an AI-generated file, and distributor screening does not grade performances. Undetectr works on Suno v6 and removes the AI watermarks and spectral fingerprints that DistroKid, Spotify and Apple Music scan for. $39 one-time, roughly 90 seconds per track.