AURA: Unified Multimodal Framework for Conversational Music Editing

arXiv GitHub Dataset

Overview

AURA is a conversational music-editing agent that listens to a song and a natural-language instruction, replies conversationally, and renders the edited audio. The system consists of three components:

  • Thinker β€” A Qwen2.5-Omni-7B model fine-tuned with LoRA (r=16, alpha=32). It processes audio and text, generates conversational replies, and emits typed edit-token blocks [EDIT_<KIND>][EDIT_0..7] (7 kinds: ADD / REMOVE / REPLACE / EXTRACT / REBALANCE / EFFECT / MOOD).
  • Bridge β€” A dual-stream fusion MusicGen decoder based on facebook/musicgen-medium. The 9 hidden states at the edit tokens condition the bridge via BiFAM (Bi-FiLM Attention Module: shared-query dual attention + FiLM modulation) and cross-attention K/V with LoRA (r=64, alpha=128). Includes learned projectors (258 MB) mapping from the thinker's hidden dimension to MusicGen's space.
  • Classifier β€” An EditSemanticClassifier (two-head: edit kind + instrument) used by the programmatic planner for stem routing.

Localized edits are code-anchored outside the requested segment and seam-crossfaded via a stem-hybrid executor (HTDemucs-6s separation).

For the dataset, this open-weight is trained on larger scale data compared to our private one in the paper to ensure the better music quality. The private weight is used for the publications, so we need to ensure we have the fair comparison.

Model Checkpoints

File Description Size
config.yaml Training configuration (paths, hyperparameters) 1 KB
thinker/adapter_config.json Thinker LoRA configuration 1 KB
thinker/adapter_model.safetensors Thinker LoRA weights (Qwen2.5-Omni-7B, r=16) 2.0 GB
bridge/projectors.pt Learned projectors (d_llm=3584 β†’ d_musicgen=2048) + FiLM MLPs/alphas/gates 258 MB
bridge/lora/adapter_config.json Bridge LoRA configuration 1 KB
bridge/lora/adapter_model.safetensors Bridge LoRA weights (MusicGen encoder_attn k/v, r=64) 37 MB
classifier/classifier.pt EditSemanticClassifier (kind + instrument heads) 14 MB

Total checkpoint size: ~2.3 GB (adapters only β€” base models downloaded separately)

Thinker Details

  • Base model: Qwen/Qwen2.5-Omni-7B
  • LoRA config: r=16, alpha=32, dropout=0.05
  • Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Modules to save: embed_tokens, lm_head (for custom edit tokens)
  • Training: 2-epoch SFT on dialogue data, then joint training with bridge (4000 steps, Ξ»=0.5 convex loss)

Bridge Details

  • Base model: facebook/musicgen-medium (1.5B params, frozen decoder)
  • Fusion mechanism: BiFAM β€” shared-query dual cross-attention over edit-token hidden states + gated FiLM modulation at each decoder layer
  • LoRA config: r=64, alpha=128, dropout=0.05, targeting encoder_attn.{k_proj, v_proj}
  • Projectors: Linear projections from thinker hidden dim (3584) to MusicGen dim (2048), plus FiLM MLP layers
  • Cross-attention layers: [0, 2, 4, 6, 8, 10, 12, 14]
  • Training: 40k steps bridge-only, then 4000 steps joint with thinker

Classifier Details

  • Architecture: Two-head classifier (edit kind: 7 classes, instrument: multi-label)
  • Input: 9 edit-token hidden states (pooled)
  • Used by: Programmatic planner for stem routing decisions

Quick Start

1. Download base models

The base models are downloaded automatically on first use, or you can pre-download them:

from huggingface_hub import snapshot_download

# Thinker base model (~15 GB)
snapshot_download("Qwen/Qwen2.5-Omni-7B", cache_dir="weights")

# Bridge base model (~3.3 GB)
snapshot_download("facebook/musicgen-medium", cache_dir="weights")

2. Download AURA checkpoints

from huggingface_hub import snapshot_download

# Download all AURA adapters (~2.3 GB)
repo_dir = snapshot_download("OpenRB-Lab/AURA")

Or download individual components:

from huggingface_hub import hf_hub_download

# Thinker LoRA adapter
thinker_config = hf_hub_download("OpenRB-Lab/AURA", "thinker/adapter_config.json")
thinker_weights = hf_hub_download("OpenRB-Lab/AURA", "thinker/adapter_model.safetensors")

# Bridge projectors + LoRA
projectors = hf_hub_download("OpenRB-Lab/AURA", "bridge/projectors.pt")
bridge_config = hf_hub_download("OpenRB-Lab/AURA", "bridge/lora/adapter_config.json")
bridge_weights = hf_hub_download("OpenRB-Lab/AURA", "bridge/lora/adapter_model.safetensors")

# Classifier
classifier = hf_hub_download("OpenRB-Lab/AURA", "classifier/classifier.pt")

3. Usage

# Point environment to your checkpoint directory
import os
os.environ["AURA_QWEN"] = "path/to/aura1/thinker"
os.environ["AURA_MG"] = "path/to/aura1/bridge"
os.environ["AURA_CLASSIFIER"] = "path/to/aura1/classifier/classifier.pt"

# Load the engine
from serving.engine import AuraEngine

engine = AuraEngine(device="cuda")
result = engine.edit(
    audio_path="path/to/song.wav",
    instruction="Add a jazzy saxophone melody to the chorus",
    guidance=2.0,
    seed=1234,
    max_seconds=5.0
)

# result["reply"]  -> conversational text response
# result["wav"]    -> edited audio (float32 numpy, 32 kHz)
# result["sr"]     -> 32000

4. Serving

# HTTP API
API_GPU=0 API_PORT=9004 bash src/scripts/serve_musicgen_api.sh

# Gradio web UI
WORKER_URL=http://127.0.0.1:9004 \
SFT_ADAPTER=path/to/aura1/thinker \
WEBAPP_PORT=7862 CUDA_VISIBLE_DEVICES=1 \
python src/edit_agent/webapp.py

Training Configuration

Training uses a 3-stage pipeline:

  1. Stage 1 β€” Thinker SFT: LoRA fine-tuning on conversational music-edit dialogues (2 epochs, lr=1e-4)
  2. Stage 2 β€” Bridge: Fusion adapter training on cached thinker hidden states (40k steps, lr=5e-5)
  3. Stage 3 β€” Joint: End-to-end training with live thinker + bridge (4k steps, loss = λ·CE_musicgen + (1βˆ’Ξ»)Β·CE_LM)

See config.yaml for the full training configuration.

Results

Production checkpoints (joint_fusion_r64/final):

Benchmark FAD ↓ CLAP ↑ SSIM ↑
IMPG Add 1.49 β€” β€”
IMPG Remove 1.36 β€” β€”
IMPG Extract 6.13 β€” β€”
Mixed 60-clip 2.15 0.661 (MuLan cos) β€”

Fusion Ablation (IMPG Benchmark)

Method FAD ↓ CLAP ↑ SSIM ↑
BiFAM (Ours) 0.41 0.218 0.776
Cross-Attention Only 2.48 0.263 0.099
Concatenation Only 13.64 0.181 0.115

Citation

@misc{trinh2026auraunifiedmultimodalframework,
      title={AURA: Unified Multimodal Framework for Conversational Music Editing}, 
      author={Quoc-Huy Trinh and Minh-Van Nguyen and Debesh Jha},
      year={2026},
      eprint={2609.14344},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2609.14344}, 
}

License

This project is licensed under the Apache License 2.0. See LICENSE for details.

Note: The base models have their own licenses:

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train OpenRB-Lab/AURA

Paper for OpenRB-Lab/AURA