SSRFNet: Stage-Wise SV-Mixer and ReDimNet Fusion Network for Speaker Verification[link]
Kyo-Won Koo, Jungwoo Heo, Seung-Bin Kim, Hyun-Seo Shin, Chan-Yeong Lim, Jisoo Son, Kyung Wha Kim, Ha-Jin Yu
abstract
Speaker verification critically depends on effective speech representations. While spectrogram-based features have long dominated, recent studies show that combining them with raw waveform signals or embeddings from large self-supervised pre-trained models (PTMs) yields complementary information. However, existing dual-branch or attention-based fusion approaches often suffer from excessive complexity or limited integration capacity. In this work, we propose the Stage-wise SV-Mixer and ReDimNet Fusion Network (SSRFNet), which retains spectrograms as the primary input and progressively injects PTM hidden states as auxiliary features via a lightweight Projection-based Fusion Module (PFM). This design ensures compatibility between heterogeneous features without requiring a full-scale additional backend. By incorporating ReDimNet as an efficient spectrogram backbone and SV-Mixer as a compressed PTM, SSRFNet achieves both high accuracy and reduced model size, with experiments confirming superior verification performance and substantially improved practicality on VoxCeleb, VoxSRC, and VCMix benchmarks.
SSRFNet: Stage-Wise SV-Mixer and ReDimNet Fusion Network for Speaker Verification[link]
Kyo-Won Koo, Jungwoo Heo, Seung-Bin Kim, Hyun-Seo Shin, Chan-Yeong Lim, Jisoo Son, Kyung Wha Kim, Ha-Jin Yu
abstract
Speaker verification critically depends on effective speech representations. While spectrogram-based features have long dominated, recent studies show that combining them with raw waveform signals or embeddings from large self-supervised pre-trained models (PTMs) yields complementary information. However, existing dual-branch or attention-based fusion approaches often suffer from excessive complexity or limited integration capacity. In this work, we propose the Stage-wise SV-Mixer and ReDimNet Fusion Network (SSRFNet), which retains spectrograms as the primary input and progressively injects PTM hidden states as auxiliary features via a lightweight Projection-based Fusion Module (PFM). This design ensures compatibility between heterogeneous features without requiring a full-scale additional backend. By incorporating ReDimNet as an efficient spectrogram backbone and SV-Mixer as a compressed PTM, SSRFNet achieves both high accuracy and reduced model size, with experiments confirming superior verification performance and substantially improved practicality on VoxCeleb, VoxSRC, and VCMix benchmarks.