Deep Learning-Based Multi-Frame Filtering for Binaural Speech Enhancement

Deep Learning-Based Multi-Frame Filtering for Binaural Speech Enhancement

Deep Learning-Based Multi-Frame Filtering for Binaural Speech Enhancement

Marvin Tammen, Simon Doclo

In many speech communication scenarios, head-mounted assistive listening devices capture not only the target speaker but also interfering sound sources, resulting in a degradation of speech quality and speech intelligibility. To alleviate this issue, several binaural speech enhancement algorithms such as the binaural minimum variance distortionless response (MVDR) beamformer have been proposed, which exploit spatial correlations of both the target speech and noise components. Furthermore, for single-microphone scenarios it has been proposed to exploit the fact that speech is highly correlated over time, resulting in the multi-frame MVDR (MFMVDR) filter. In this contribution, we consider a binaural extension of the MFMVDR filter, which exploits both spatial as well as temporal correlations. The binaural MFMVDR filter is embedded into an end-to-end deep learning framework, where the required parameters are estimated by temporal convolutional networks (TCNs) that are trained by minimizing the mean spectral absolute error loss function. Simulation results comprising measured binaural room impulses and diverse noise sources at signal-to-noise ratios in [-5, 20] dB demonstrate the ad- vantage of utilizing the binaural MFMVDR filter structure instead of directly estimating the multi-frame filter coefficients with TCNs.

Algorithm Bagpipe, 3dB Frogs, 19dB Munching, 7dB Fan, 4dB
noisy
clean
no imposed structure
imposed binaural MFMVDR structure
(Changed: 24 Jun 2026)  Kurz-URL:Shortlink: https://uol.de/p92367en
Zum Seitananfang scrollen Scroll to the top of the page

This page contains automatically translated content.