ERASYOpen App

AI Music

The AI Sound Signature: What You Hear vs What Machines Measure

By Eddie Mathews··11 min read
The AI Sound Signature: What You Hear vs What Machines Measure — Erasy

Ask in any AI music community whether there is an AI sound signature and you will get two confident, incompatible answers. Producers describe it precisely — the glassy sustain, the smeared consonants, the washed reverb tail — and say they hear it within seconds. Meanwhile a blind test of 9,000 people across eight countries found that 97% could not pick a fully AI-generated track out of a lineup at all.

Both findings are real, and the reason they do not contradict each other is the most useful thing anybody can tell you about your own tracks. There is not one AI signature. There are three, they have different causes, and fixing one does nothing whatsoever for the other two. We went through the Deezer and Ipsos survey data, the ISMIR 2025 paper that explains the machine-detectable artifact, and Deezer's own production detection figures to separate them properly.

Key takeaways

97% of listeners cannot identify AI music. Deezer and Ipsos, 9,000 adults across eight countries, October 2025 — a blind test of two AI tracks and one human one.

A production detector gets it right 99.8% of the time. Deezer, July 2026: about two misses per 1,000 AI tracks, under one false positive per 10,000 human ones.

The gap is the whole story. What people hear and what machines measure are different phenomena. The detectable one is inaudible by nature.

Most of what you can hear is an unfinished file. Raw generator exports are unmastered and uncorrected, and that is a large part of what reads as synthetic.

Suno V6 moves one layer, probably not the others. The research says the detectable artifact comes from architecture, not from training data — and training data is what V6 changed.

The 97% problem, and why it is not a contradiction

In October 2025 Deezer commissioned Ipsos to run what it described as the first survey of public attitudes to AI-generated music. The sample was 9,000 adults aged 18 to 65 across the United States, Canada, Brazil, the UK, France, the Netherlands, Germany and Japan. Part of it was a blind test: two fully AI-generated songs and one human-made track. 97% could not tell which was which, and 52% said they were uncomfortable about that.

Set that next to Deezer's own detection numbers from July 2026 and the shape of the problem appears. The same company that found humans at 3% runs an automated detector it puts at 99.8% accuracy, across roughly 90,000 fully AI-generated tracks arriving every day — more than half of everything uploaded, at the June 2026 peak.

Human listenersDeezer's detector
Correct identification3%99.8%
What it got wrong97% of the time~2 misses per 1,000 AI tracks
False positivesNot measured<1 per 10,000 human tracks
Sample9,000 adults, 8 countriesEvery upload, ~90,000/day
ConditionsFinished, mastered tracksSignal-level analysis
DateOctober 2025July 2026

A thirty-point gap would be a disagreement. A gap this size means the two are not measuring the same thing at all. The listeners heard finished records once, the way anybody hears music. The detector reads structure in the signal that is not available to hearing in the first place. So when a producer insists they can hear it and a survey says almost nobody can, both can be telling the truth — because the producer is describing something else entirely.

Three different things wearing one name

Almost every guide to spotting AI music treats the subject as a single checklist. That is the error worth fixing, because the three phenomena people bundle together have separate causes, separate detectability and separate remedies. Getting better at one of them is not progress on the others, and a lot of wasted effort comes from not knowing which one you are actually working on.

LayerWhat it isAudible?What catches itThe fix
1 — Decode artifactsGlassy sustains, smeared consonants, washed reverb tailsSometimes, on raw exportsAttentive ears; not detectorsRepair and mastering
2 — Spectral peaksSmall regular frequency peaks from the model's architectureNo — neverAutomated screeningSignal-level processing before submission
3 — Compositional defaultSafe progressions, conventional arrangement, tidy structureYes, easilyAny attentive listenerPrompting, editing, arrangement

Layer three is what most people mean when they say a track sounds like AI, and it is the layer with the least to do with the technology. Layer two is the only one that decides whether your release goes live. Layer one sits in between, and it is the one that has changed most with each model generation.

Layer one: what you are actually hearing

Current music generators do not paint a waveform directly. They work in a compressed representation and then hand it to a decoder that turns it back into audio. That compression step is lossy by design: the model rounds a continuous internal representation onto a fixed set of learned values, and the difference between the two is not recovered on the way out. What arrives in your export is an approximation, reconstructed.

This is why the complaints cluster where they do. Sustained material suffers most, because a held note gives the reconstruction the longest stretch with nothing new to anchor to — hence the glassiness on long vowels. Transients suffer next, because a sharp attack is precisely the thing a compressed representation handles least gracefully, which is where the smeared consonants and slightly soft drum hits come from. Reverb tails are the third, for the same reason as sustains.

The honest caveat

The machine-detectable artifact in layer two has a peer-reviewed explanation. The audible layer described here does not.

What exists is a consistent body of listener reports that name the same handful of characteristics, and a decoder architecture whose known behaviour is a plausible cause. That is a reasonable inference and we are labelling it as one. Anybody presenting the audible AI sound as a measured, published phenomenon is going further than the evidence.

The practical point is that this layer responds to ordinary studio work. Corrective processing, de-essing, careful level discipline and a proper loudness target address most of it, which is why the same track that screamed synthetic as a raw export can sit unremarkably in a playlist after mastering. A large share of what people identify as the AI sound is really the sound of a file nobody has finished yet.

Layer two: the signature you cannot hear

The artifact that actually governs your release has a proper explanation, and it won best paper at ISMIR 2025. Darius Afchar, Gabriel Meseguer-Brocal, Kamil Akesbi and Romain Hennequin of Deezer published A Fourier Explanation of AI-music Artifacts, showing that the deconvolution modules common to generative audio models produce systematic frequency artifacts — small but distinctive spectral peaks, related to the checkerboard artifact familiar from image generation.

99%+
Detection accuracy from the peaks alone
Architecture
Not training data or weights
Suno & Udio
Among the generators validated
0%
Of it is audible to you

Two findings in that paper matter more than the headline. The first is that detection built on these peaks alone reaches accuracy on par with deep-learning approaches, surpassing 99% in several scenarios — no trained classifier required, just the physics of the operation. The second is the sharper one: the peaks are inherent to the chosen model architecture rather than a consequence of the training data or the model weights.

That second finding is why this layer behaves so differently from the other two. It is not a watermark somebody chose to add and could choose to remove. It is a by-product of how the audio gets built, present whether or not the generator intended to mark anything, and it does not care how good your track sounds. It is also why two detectors can disagree about the same file while both being broadly accurate: they are not all reading the same features.

Layer three: sameness is not an artifact

The third layer is the one listeners complain about most and the one with no technical cause at all. Hand a model a broad genre tag and an emotional adjective and it settles toward the centre of what it has learned, because that is what the objective rewards. The result is the progression you have heard a thousand times, an arrangement that resolves exactly where you expect, and a vocal performance that takes no risks.

Reads as generated
  • Broad genre tag plus an emotional adjective
  • First or second generation accepted as final
  • No breath sounds, or breaths in places no singer would take them
  • Structure that resolves exactly where expected
  • Every section at the same energy
Reads as a record
  • Specific instrumentation, era and production references
  • Many generations, edited and stitched
  • Arrangement decisions the model did not make for you
  • One deliberate departure from the obvious
  • Dynamics across the track, not just within a bar

This is worth stating plainly because it is the layer entirely within your control, and it is the reason two people using the identical model produce work that sounds nothing alike. It is also the layer that no amount of processing will help. A generic song cleaned to broadcast spec is a clean generic song.

Does Suno V6 remove the AI sound signature?

This is the question that prompted the article — it surfaced on r/SunoAI as someone asking whether they had heard the characteristic AI signature in the V6 teaser or were imagining it. V6 shipped on 9 September 2026 as a family of three models built with Warner Music Group, BMG and Believe, and a jump in audio quality is the headline claim. Our day-two review covers what else changed.

Answering it properly means answering it three times, because the three layers move independently.

LayerDid V6 move it?Reasoning
1 — Audible decode artifactsProbably improvedFidelity is the headline claim of the release, and this is the layer fidelity work lands on
2 — Spectral peaksNo reason to expect a changeThe ISMIR paper attributes them to architecture, not to training data or weights — and licensed data is what V6 changed
3 — Compositional defaultMixed, and by designThe flagship is more literal about following direction; v6-wild exists specifically to restore looser, more characterful output

The middle row is the one to sit with. If someone tells you a new model generation has beaten AI detection, ask what they think changed. A better decoder that is still a decoder still deconvolves. To clear layer two, the architecture itself would have to change in the specific way the paper identifies, and a fidelity improvement is not that.

Nobody has published a measurement of V6 against the ISMIR method, so this is reasoning from the paper rather than a tested result. We would rather say that than give you a confident number we do not have.

Test your own track in ten minutes

You cannot audit layer two by ear, and you should stop trying. What you can do is find out which of the other two layers your track is losing on, which takes about ten minutes and settles most arguments with yourself.

01
Level-match against a reference

Pick a commercial track in your genre and match loudness before you compare anything. Almost every 'sounds AI' judgement made at mismatched levels is really a judgement about loudness.

02
Listen to the vocal alone

Solo it and listen to sustained notes and consonants specifically. This is where layer one concentrates, and where it either is or is not a problem on your material.

03
Check the breaths

Are there any, and are they where a singer would take them? Absent or misplaced breathing is the single most reported giveaway, and it is fixable in an edit.

04
Write down the chord progression

If you can predict the next chord every time on first listen, that is layer three, and no processing will touch it. Go back to the generation stage.

05
Master it, then repeat

Run the same comparison after a proper master. Whatever survives is a genuine layer-one problem; whatever disappeared was never the AI, it was the unfinished file.

06
Accept what you cannot test

Layer two is not available to this process at any level of skill or monitoring. Treat it as a separate step in the release workflow rather than a listening problem.

What this actually means when you release

Here is where the three layers stop being an interesting distinction and start costing money. Layers one and three decide whether people enjoy the track. Layer two decides whether it gets in front of them at all — distributors run automated screening at the moment you submit, and a flagged upload does not come back with a note explaining itself. It simply does not go live.

Worth being precise about what that screening is and is not. It is not a ban: DistroKid, RouteNote, UnitedMasters, LANDR, Amuse and Symphonic all accept AI music openly. And clearing it has nothing to do with platform labelling — Spotify's AI Persona badge and Apple's transparency tags are disclosure systems, declared on delivery, and no tool can or should interfere with them. The screening problem is narrower and more mechanical than either: a signal-level check your file either passes or does not.

LayerWhat it decidesWhere it gets handled
1 — Audible artifactsWhether listeners stay past the first chorusRepair and mastering, before you bounce
2 — Spectral peaksWhether the release goes live at allSignal-level cleanup, after the master, before you submit
3 — CompositionWhether anyone plays it twiceThe generation stage — nothing downstream fixes it

That is the gap Undetectr is built for — it targets the watermark signals and generation artifacts distributors screen for, at the signal level where layer two lives rather than the audible level where mastering works. It does not make your track stop being AI-generated and it cannot affect a platform's disclosure label. Our 12-step release checklist puts that step in order with the rest.

The layer you cannot hear

Mastering fixes what you hear. Screening reads what you cannot.

Undetectr targets the watermark signals and generation artifacts distributors screen for, so your track clears intake at Spotify, Apple Music and beyond. It will not change how a platform labels your release — nothing can. €39 one-time, no subscription.

And then the harder problem

Clearing screening gets you distributed, which is worth being honest about: distribution is the solved part. Deezer counts AI music at 1–3% of actual listening against more than half of daily uploads, so the wall is not getting in, it is being heard once you are. If you want somewhere that does not depend on algorithmic discovery, pitching for paid sync placements in TV, film, games and ads is where the money conversation in this niche actually is — and selling direct to the listeners you already have keeps 100% of something rather than a fraction of very little.

Frequently asked questions

Erasy

Eddie Mathews — AI Music Editor, Erasy

Eddie covers the craft side of AI music for Erasy — prompting, arrangement, vocal direction and the workflow between a generation you like and a file you can release. They test on their own material and say plainly where a technique stops working.

Is the AI sound signature something you can actually hear?+
Partly, and the part you can hear is not the part that matters for screening. Producers listening closely to raw generator exports consistently report the same handful of things: a glassy sheen on sustained vocals, consonants that smear rather than snap, reverb tails that sound washed rather than decayed, and a slightly flat stereo image. Those are real and they are audible. The signature that distributors and streaming platforms actually screen for is a set of small spectral peaks that no human ear resolves at all. The two are different phenomena with different causes, and improving one does nothing for the other.
If AI music has a signature, why did 97% of listeners fail to spot it?+
Because the listeners in that test were not doing the thing producers do. Deezer commissioned Ipsos to run a blind test with 9,000 adults across eight countries in October 2025, playing two fully AI-generated songs and one human-made track, and 97% could not reliably tell them apart. Those were finished, mastered tracks heard once, in the way people normally hear music. A producer A/B-ing their own raw export against a commercial reference, on monitors, at the moment of maximum attention, is running an entirely different test. Both results are honest; they are measuring different things.
Does mastering remove the AI sound signature?+
It removes a good deal of the audible layer and none of the machine-detectable one. Much of what people describe as the AI sound in a raw export is really the sound of an unfinished file: no corrective processing, no level discipline, no loudness target. Mastering addresses that, which is why cleaned-up AI tracks stop sounding obviously synthetic to casual listeners. It does not touch the spectral peaks a detector reads, because those sit at a level of the signal that mastering is not aiming at and is not designed to alter.
Does Suno V6 remove the AI sound signature?+
It plausibly improves the audible layer and there is no reason to expect it removed the detectable one. V6 shipped on 9 September 2026 built with Warner Music Group, BMG and Believe, and a jump in audio quality is Suno's headline claim, so the glassiness and smear people complain about should be less pronounced. The detectable signature is a different matter. The ISMIR 2025 paper that explains it found the spectral peaks are inherent to the model architecture rather than to the training data or the model weights — and licensed training data is precisely what changed with V6. Nobody has published a measurement either way, so treat this as reasoning from the paper, not as a tested result.
Can a detector be wrong about my track?+
Yes, in both directions, though the published error rates are small. Deezer reports its production detector running at 99.8% accuracy, which it breaks down as roughly two missed detections per 1,000 AI tracks and fewer than one false positive per 10,000 human-made tracks. Those are good numbers, but at the scale of a streaming catalogue they still mean real human records get flagged and real AI records slip through. A detector returns a probability, not a ruling, and different detectors disagree with each other on the same file more often than their individual accuracy figures suggest.
Why do my AI vocals sound thin or distant compared to the instruments?+
Vocals are where the audible layer concentrates, and there is a straightforward reason. A voice is the sound listeners have the most exposure to and the strongest expectations about, so small departures from a natural performance register as wrong far faster than the same departure in a pad or a hi-hat. Breath placement is the usual giveaway: absent entirely, or in a spot no singer would take one. Sustained notes are the second, where the tail of a held vowel is the hardest thing for a generator to keep convincing. Neither is a watermark and neither is what a scanner is reading.
Is the generic-sounding quality of AI music part of the signature?+
It is the part people notice most and the part that is least to do with the technology. Given a broad genre tag and an emotional adjective, a model settles toward the centre of what it has learned, which means common progressions, conventional arrangements and unsurprising vocal choices. That is a prompting and editing outcome, not an artifact of generation, and it is the one layer entirely within your control. It is also the reason two people using the same model can produce work that sounds nothing alike.

Disclosure: Erasy is an independent guide to AI music cleanup, and Undetectr is the tool we recommend and link to. Survey and detection figures are quoted from Deezer's own newsroom; the spectral-artifact findings are from Afchar et al., "A Fourier Explanation of AI-music Artifacts", ISMIR 2025. Where this article reasons from that paper to Suno V6 rather than citing a measurement, it says so in the text. No published test of V6 against that method exists as of writing. Figures current as of 11 September 2026.