Publication

해외 컨퍼런스SSRFNet: Stage-Wise SV-Mixer and ReDimNet Fusion Network for Speaker Verification

88f9ff1823a54.png


SSRFNet: Stage-Wise SV-Mixer and ReDimNet Fusion Network for Speaker Verification​[link]​​​


Kyo-Won Koo, Jungwoo Heo, Seung-Bin Kim, Hyun-Seo Shin, Chan-Yeong Lim, Jisoo Son, Kyung Wha Kim, Ha-Jin Yu


abstract


Speaker verification critically depends on effective speech representations. While spectrogram-based features have long dominated, recent studies show that combining them with raw waveform signals or embeddings from large self-supervised pre-trained models (PTMs) yields complementary information. However, existing dual-branch or attention-based fusion approaches often suffer from excessive complexity or limited integration capacity. In this work, we propose the Stage-wise SV-Mixer and ReDimNet Fusion Network (SSRFNet), which retains spectrograms as the primary input and progressively injects PTM hidden states as auxiliary features via a lightweight Projection-based Fusion Module (PFM). This design ensures compatibility between heterogeneous features without requiring a full-scale additional backend. By incorporating ReDimNet as an efficient spectrogram backbone and SV-Mixer as a compressed PTM, SSRFNet achieves both high accuracy and reduced model size, with experiments confirming superior verification performance and substantially improved practicality on VoxCeleb, VoxSRC, and VCMix benchmarks.


본사이트의 모든 제작물의 저작권은 IRLab에 있으며, 무단복제나 도용은 저작권법(96조)에 의해 금지되어 있습니다.

COPYRIGHT ©  IRLab . Ltd. ALL RIGHTS RESERVED.