UFL-GAN: A Multi-Discriminator GAN for Unsupervised Speech Enhancement

Satvik Bejugam, Venkatesh Parvathala, K. Sri Rama Murty
Abstract: Most deep learning based speech enhancement methods are usually trained in a supervised fashion, i.e., they typically rely on parallel corpora of noisy and clean speech pairs. This is often difficult to obtain in real-world scenarios, leading to the use of synthetic data. In this work, we propose a novel unsupervised speech enhancement method that does not require paired training data. We introduce a multi-discriminator GAN-based architecture to capture global (or utterance-level) and local (or frame-level) characteristics of the speech signal. Additionally, we incorporate self-supervised representations from a pre-trained model to provide auxiliary information to the generator and thereby enhance the denoising capability. Our extensive experimental results on the VoiceBank+DEMAND dataset demonstrate that the proposed method achieves comparable or better performance across both intrusive and non-intrusive quality measures.
UFL-GAN model architecture
Architecture of the proposed UFL-GAN.
Each example shows a spectrogram montage (Clean / Noisy / U / F / UF) above the corresponding audio. U uses an utterance-level discriminator, F a frame-level discriminator, and UF, the proposed approach, combines both.

1. p232_328.wav

Grayscale spectrogram montage for p232_328.wav
Clean
Noisy - 1.22
U (utterance) - 1.47
F (frame) - 1.23
UF (proposed) - 1.62

2. p232_160.wav

Grayscale spectrogram montage for p232_160.wav
Clean
Noisy - 1.09
U (utterance) - 1.15
F (frame) - 1.11
UF (proposed) - 1.19

3. p232_203.wav

Grayscale spectrogram montage for p232_203.wav
Clean
Noisy - 1.11
U (utterance) - 1.20
F (frame) - 1.13
UF (proposed) - 1.25