
The audio task
Anyone stocking a catalogue, a sample library, or a playlist with AI-generated tracks across several genres is implicitly betting that a generation system produces the same range of variation a human catalogue would, an assumption a computational audit can test rather than assume.
What the documents show
A preprint titled The Algorithmic Flattening of Sound: Computational Evidence and Justice Implications of AI Music Homogenization, posted to arXiv on 6 August 2026 by Zoe Slendebroek and Danaé Metaxa, generated 100 tracks per genre from two commercial systems, Suno (Pro tier, version 5.5) and Lyria 3, across Afrobeats, K-pop, dance pop, and heavy metal, using feature-matched and minimal genre-name prompts collected in April and May 2026. Comparing 72 music-information-retrieval features against human reference tracks, the full text reports that Lyria 3's output shows reduced acoustic diversity within a genre, while Suno's output instead collapses the acoustic distinctions between genres without similarly narrowing what varies inside each one. A classifier trained only on those MIR features separated AI from human tracks with a mean AUC of 0.991 and accuracy of 0.967, a separation the paper says held up even when restricted to instrumental tracks, which it interprets as evidence the gap is not simply about vocals. The authors frame the pattern as a cultural, economic, and epistemic justice question: because generation systems reproduce whatever is statistically dominant in their training data, they argue non-Western genres risk being flattened toward under-represented, unrepresentative samples.
Rights status
This is a technical audit, not a legal filing; it makes no copyright or licensing claim about either system's training data or output, and none should be read into it. Its evidentiary weight concerns measured acoustic similarity and classifier separability, not the legal status of AI-generated music.
What to check before you use it
This is an editorial checklist. The paper is a preprint, not yet peer-reviewed at the point of this retrieval, and the authors state plainly that their findings do not estimate how common this pattern is across other models, genres, or time periods beyond the two systems and four genres tested. Confirm the exact system versions, Suno Pro tier 5.5 and Lyria 3, before extending the finding to a newer release, since the paper says it characterises output under realistic use conditions rather than the best case a more elaborate prompt might achieve. Note the human comparison corpus was limited to 30-second preview clips, a constraint the authors flag as potentially missing variation present in full tracks.
- Does a claim apply to the specific model versions tested, or is it being generalised to AI music broadly?
- Is the classifier's near-perfect accuracy being read as proof any listener, rather than a trained model on MIR features, can hear the difference?
- Has this preprint been revised or peer-reviewed since the 6 August 2026 posting cited here?
Read narrowly, the preprint's contribution is a method and a named-model finding, two systems flattening sound in two different directions, rather than a general verdict on generative music, and any reuse of its numbers should keep that scope attached.
Sources & reading trail
Abstract gives the tested systems, genres, and top-line homogenization and classifier findings.
Source published: 6 August 2026 · Retrieved: 16 September 2026
Full text gives the prompting method, exact classifier accuracy figures, and the paper's own stated scope limitations.
Source published: Not established · Retrieved: 16 September 2026
Documentation, licences and platform policies establish the note; the what-to-check reading is Music Tech Field Notes editorial analysis. This retrospective draft does not imply the site published on the event date.