CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
A paper titled 'CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation' was published on arXiv on 2026-09-11. It proposes a method called Curriculum-guided Multi-scale representation Alignment with scale-aware Optimal Transport (CMA-OT) that uses an external music expert to provide hierarchical supervision for a music generator's latent features, aiming to improve dance-to-music generation.
The paper addresses dance-to-music (D2M) generation, which synthesizes music aligned with dance videos. It identifies a semantic mismatch between sparse dance cues (rhythm, style) and dense music composition requirements. Existing methods supervise only final audio output, leading to poor music representations. CMA-OT introduces hierarchical supervision from an external music expert and a curriculum-guided multi-scale learning strategy to progressively transfer musical knowledge, enhancing representation learning and generated music quality.
The approach uses scale-aware optimal transport to align multi-scale representations between the expert and generator, with curriculum guidance for stable training. This suggests a shift toward leveraging pre-trained music models as teachers for generative tasks, potentially improving structural coherence and musicality.
This research could influence AI music generation tools by enabling more context-aware and stylistically consistent music synthesis from visual inputs, relevant to creative industries, gaming, and social media content creation.
Improved dance-to-music generation could enhance automated content creation pipelines, reducing manual music composition costs for video producers and enabling new interactive entertainment experiences.
Next signals include follow-up papers applying similar hierarchical supervision to other cross-modal generation tasks, or integration into commercial music generation platforms. Watch for benchmarks comparing CMA-OT against existing D2M methods.