Papers
arxiv:2609.33757

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Published on Sep 27
· Submitted by
Ruibin Yuan
on Sep 29
#3 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.

Community

We're sharing the YuE2 technical report. YuE2 unifies symbolic and audio music generation: it first writes an editable melody-and-chord score, then renders a full song with vocals and accompaniment. The same model supports zero-shot covers and score-based editing, including music editing guided by an external language-model agent. The report also introduces MERT2 and SheetSage2 for music understanding and transcription, and includes automatic benchmarks and expert listening evaluations.

Model weights: https://huggingface.co/m-a-p/YuE2-3B
Demos: https://map-yue2.github.io/
Code: https://github.com/multimodal-art-projection/YuE
PDF: https://github.com/multimodal-art-projection/YuE/blob/main/docs/technical_report.pdf

Paper submitter
•
This comment has been hidden (marked as Resolved)

YUE2 reprents deepseek moment for new music composition with open model.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33757
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 6

Browse 6 models citing this paper

Datasets citing this paper 1

Spaces citing this paper 59

Browse 59 spaces citing this paper

Collections including this paper 2