Artificial Analysis Launches AA-Music v1.1: Suno v6 Sweeps Top Ranks
Artificial Analysis released AA-Music v1.1 across 17 genres and 1,000 prompts with 37,000 blind panel votes. Suno v6 took first place in both vocal and instrume
On October 8, 2026, independent AI benchmarking platform Artificial Analysis officially rolled out its updated music generation benchmarks, 'AA-Music-Vocal v1.1' and 'AA-Music-Instrumental v1.1'. Built on 1,000 newly curated prompts across 17 musical genres and scored using more than 37,000 blind human preference evaluations, Suno's flagship model Suno v6 captured the undisputed number one spot across both vocal and instrumental categories.

Image source: Artificial Analysis (@ArtificialAnlys)
The v1.1 update powers the Artificial Analysis Music Arena and separates evaluation into two standalone leaderboards: vocal tracks featuring sung or spoken vocal elements, and pure instrumental pieces. Over the past three months, flagship model releases have elevated baseline AI music quality from short loops into fully arranged songs with authentic instrumentation, making nuanced, rigorous human evaluation increasingly critical to discern differences between top systems.
1,000 Prompts Across 17 Genres and 37,000 Blind Panel Votes
The core technical enhancement in AA-Music v1.1 is the introduction of an expanded, genre-balanced test set that mirrors real-world prompting behavior.
The test suite comprises 1,000 new prompts—500 allocated to the vocal benchmark and 500 to the instrumental benchmark—curated across 17 distinct musical genres. The coverage spans classical and jazz & blues to electronic dance music (EDM), heavy metal, reggae, pop, hip-hop, Memphis soul, and rocksteady. Rather than simple keyword tags, each prompt specifies intricate song structure, emotional mood, instrument arrangements, and vocal styling to thoroughly test the upper boundaries of frontier music models.
- Strictly Evaluator Panel Blind Votes: Official leaderboard ratings are derived exclusively from blind preference votes cast by Artificial Analysis's recruited human evaluator panel. Public Music Arena votes do not count toward official ratings, ensuring controlled and consistent scoring standards.
- Statistical Rigor and Sample Depth: Evaluators listened to sample pairs generated from identical prompts without model attribution and selected their preferred track. The launch encompasses over 37,000 blind human votes (more than 16,000 on vocal and 21,000 on instrumental), guaranteeing over 2,000 votes for every ranked model.
- Bradley-Terry Elo Scaling: Scores are computed via Bradley-Terry Maximum Likelihood Estimation (MLE) and scaled into an Elo-style metric alongside a 95% confidence interval, clearly delineating statistical ties from verified performance advantages.
Suno v6 Dominates Both Leaderboards Alongside Generational Gains
The initial leaderboards establish Suno v6 as the frontrunner across both modalities, securing the top two positions on each board alongside its smaller companion model.
On the AA-Music-Vocal v1.1 leaderboard, Suno v6 took first place with an Elo rating of 1142 (95% confidence interval 1127–1157). It holds a 26 Elo advantage over second-place Suno v6-mini (Elo 1116), allowing Suno to capture the top two spots. On AA-Music-Instrumental v1.1, Suno v6 maintained its lead with an Elo score of 1140, standing 31 Elo ahead of Suno v6-mini (Elo 1109).
- Mureka V9.5 Generational Leap: Mureka V9.5 secured third place on the vocal leaderboard with an Elo rating of 1103. This represents a 51 Elo leap over its predecessor Mureka V9 (Elo 1052), marking the largest generational gain between two iterations of the same model family on the vocal board.
- Instrumental Statistical Dead Heat: In the instrumental category, four competing models landed in a tight statistical tie for ranks three through six: Mureka V9 (Elo 1081), Mureka V9.5 (Elo 1080), Lyria 3 Pro (Elo 1078), and Lyria 3.5 (Elo 1076), all separated by just 5 Elo points.
- Eleven Music and Key Model Rankings: Eleven Music v2.5 placed seventh on the vocal leaderboard (Elo 1012), while base Eleven Music reached seventh on instrumental (Elo 1021). On the instrumental leaderboard, Stable Audio 3 Large placed eighth with an Elo score of 1017. Among open-weights models, MiniMax Music 3.0 emerged as the top performer, ranking 11th on vocals (fixed as the 1000 Elo reference anchor) and 13th on instrumental.
End-to-End Songwriting Evaluation and Production Workflow Trade-Offs
A key methodological cornerstone of AA-Music-Vocal v1.1 is the deliberate omission of pre-written lyrics.
By excluding user-provided lyrics from the benchmark prompts, the evaluation measures each model's end-to-end songwriting capability. Models must independently compose coherent thematic lyrics, structure musical verses and choruses, sculpt vocal phrasing, and blend instrumental backing into a cohesive track from descriptive text alone. Frontier systems have moved beyond rudimentary loops into fully orchestrated musical compositions.
For creative producers and commercial developers, however, leaderboard ratings represent only one component of model selection.
- Panel Consensus Versus Audience Taste: Because the official leaderboard relies strictly on controlled evaluator panel decisions, preferences may diverge from specific subgenre niches or viral social media listener dynamics.
- Cost and Licensing Realities: As performance gaps between leading models compress within 20 to 30 Elo points, practical adoption is increasingly dictated by API generation pricing per track and commercial licensing rights rather than incremental audio fidelity.
- Genre-Specific Variance: Individual models exhibit varying strengths across different genres, excelling in synth-heavy electronic pop while struggling with organic acoustic live-band nuances, making tailored testing essential for production use cases.
Sources
- Artificial Analysis Official X (@ArtificialAnlys): AA-Music v1.1 Benchmark Release Thread
- Artificial Analysis: AA-Music-Vocal v1.1 Leaderboard
- Artificial Analysis: AA-Music-Instrumental v1.1 Leaderboard
- Artificial Analysis: Music Generation Benchmarking Methodology