
Mu-Mixer: A Hierarchical MLP-Based Framework for Cover Song Identification[link]
Jungwoo Heo, Hyun-Seo Shin, Chan-Yeong Lim, Kyo-Won Koo, Seung-Bin Kim, Jisoo Son, Ha-Jin Yu
abstract
Cover Song Identification (CSI) is a fundamental task in music information retrieval, essential
for copyright protection and content organization. While Convolutional Neural Networks (CNNs) and Transformers have significantly advanced CSI, they suffer from limited receptive fields and quadratic computational complexity, respectively. In this paper, we propose Mu-Mixer, a novel all-MLP architecture tailored for CSI that achieves global receptive fields with linear complexity. In contrast to image-based MLP-Mixers, which indiscriminately mix information across the temporal axis, Mu-Mixer incorporates a Multi-scale Temporal Mixing Block (MTMB). This strategy hierarchically segments the time dimension T into multiple scales to preserve local rhythmic motifs while capturing long-range structural dependencies. We evaluate Mu-Mixer on the AI-Hub Music Similarity dataset, where it achieves state-of-the-art performance. Our results demonstrate that Mu-Mixer provides a superior balance between identification accuracy and computational efficiency, proving that self-attention is not a prerequisite for modeling long range acoustic dependencies in musical signal
Mu-Mixer: A Hierarchical MLP-Based Framework for Cover Song Identification[link]
Jungwoo Heo, Hyun-Seo Shin, Chan-Yeong Lim, Kyo-Won Koo, Seung-Bin Kim, Jisoo Son, Ha-Jin Yu
abstract
Cover Song Identification (CSI) is a fundamental task in music information retrieval, essential
for copyright protection and content organization. While Convolutional Neural Networks (CNNs) and Transformers have significantly advanced CSI, they suffer from limited receptive fields and quadratic computational complexity, respectively. In this paper, we propose Mu-Mixer, a novel all-MLP architecture tailored for CSI that achieves global receptive fields with linear complexity. In contrast to image-based MLP-Mixers, which indiscriminately mix information across the temporal axis, Mu-Mixer incorporates a Multi-scale Temporal Mixing Block (MTMB). This strategy hierarchically segments the time dimension T into multiple scales to preserve local rhythmic motifs while capturing long-range structural dependencies. We evaluate Mu-Mixer on the AI-Hub Music Similarity dataset, where it achieves state-of-the-art performance. Our results demonstrate that Mu-Mixer provides a superior balance between identification accuracy and computational efficiency, proving that self-attention is not a prerequisite for modeling long range acoustic dependencies in musical signal