The Experts below are selected from a list of 360 Experts worldwide ranked by ideXlab platform
Rémi Gribonval - One of the best experts on this subject based on the ideXlab platform.
-
Spatial location priors for Gaussian model based reverberant Audio Source Separation
EURASIP Journal on Advances in Signal Processing, 2013Co-Authors: Ngoc Q. K. Duong, Emmanuel Vincent, Rémi GribonvalAbstract:We consider the Gaussian framework for reverberant Audio Source Separation, where the Sources are modeled in the time-frequency domain by their short-term power spectra and their spatial covariance matrices. We propose two alternative probabilistic priors over the spatial covariance matrices which are consistent with the theory of statistical room acoustics and we derive expectation-maximization algorithms for maximum a posteriori (MAP) estimation. We argue that these algorithms provide a statistically principled solution to the permutation problem and to the risk of overfitting resulting from conventional maximum likelihood (ML) estimation. We show experimentally that in a semi-informed scenario where the Source positions and certain room characteristics are known, the MAP algorithms outperform their ML counterparts. This opens the way to rigorous statistical treatment of this family of models in other scenarios in the future.
-
A tractable framework for estimating and combining spectral Source models for Audio Source Separation
Signal Processing, 2012Co-Authors: Simon Arberet, Alexey Ozerov, Frédéric Bimbot, Rémi GribonvalAbstract:The underdetermined blind Audio Source Separation (BSS) problem is often addressed in the time-frequency (TF) domain assuming that each TF point is modeled as an independent random variable with sparse distribution. On the other hand, methods based on structured spectral model, such as the Spectral Gaussian Scaled Mixture Models (Spectral-GSMMs) or Spectral Non-negative Matrix Factorization models, perform better because they exploit the statistical diversity of Audio Source spectrograms, thus allowing to go beyond the simple sparsity assumption. However, in the case of discrete state-based models, such as Spectral-GSMMs, learning the models from the mixture can be computationally very expensive. One of the main problems is that using a classical Expectation-Maximization procedure often leads to an exponential complexity with respect to the number of Sources. In this paper, we propose a framework with a linear complexity to learn spectral Source models (including discrete state-based models) from noisy Source estimates. Moreover, this framework allows combining di erent probabilistic models that can be seen as a sort of probabilistic fusion. We illustrate that methods based on this framework can significantly improve the BSS performance compared to the state-of-the-art approaches.
-
Under-Determined Reverberant Audio Source Separation Using Local Observed Covariance and Auditory-Motivated Time-Frequency Representation
2010Co-Authors: Ngoc Duong, Emmanuel Vincent, Rémi GribonvalAbstract:We consider the local Gaussian modeling framework for under-determined convolutive Audio Source Separation, where the spatial image of each Source is modeled as a zero-mean Gaussian variable with full-rank time- and frequency- dependent covariance. We investigate two methods to improve the accuracy of parameter estimation, based on the use of local observed covariance and auditory-motivated time-frequency representation. We derive an iterative expectation-maximization (EM) algorithm with a suitable initialization scheme. Experimental results over stereo synthetic reverberant mixtures of speech show the effectiveness of the proposed methods.
-
Nonnegative matrix factorization and spatial covariance model for under-determined reverberant Audio Source Separation
2010Co-Authors: Simon Arberet, Alexey Ozerov, Frédéric Bimbot, Rémi Gribonval, Emmanuel Vincent, Ngoc Duong, Pierre VandergheynstAbstract:We address the problem of blind Audio Source Separation in the under-determined and convolutive case. The contribution of each Source to the mixture channels in the time-frequency domain is modeled by a zero-mean Gaussian random vector with a full rank covariance matrix composed of two terms: a variance which represents the spectral properties of the Source and which is modeled by a nonnegative matrix factorization (NMF) model and another full rank covariance matrix which encodes the spatial properties of the Source contribution in the mixture. We address the estimation of these parameters by maximizing the likelihood of the mixture using an expectation-maximization (EM) algorithm. Theoretical propositions are corroborated by experimental studies on stereo reverberant music mixtures.
-
Spatial covariance models for under-determined reverberant Audio Source Separation
2009Co-Authors: Ngoc Duong, Emmanuel Vincent, Rémi GribonvalAbstract:The Separation of under-determined convolutive Audio mixtures is generally addressed in the time-frequency domain where the Sources exhibit little overlap. Most previous approaches rely on the approximation of the mixing process by complex-valued multiplication in each frequency bin. This is equivalent to assuming that the spatial covariance matrix of each Source, that is the covariance of its contribution to all mixture channels, has rank 1. In this paper, we propose to represent each Source via a full-rank spatial covariance matrix instead, which better approximates reverberation. We also investigate a possible parameterization of this matrix stemming from the theory of statistical room acoustics. We illustrate the potential of the proposed approach over a stereo reverberant speech mixture.
Emmanuel Vincent - One of the best experts on this subject based on the ideXlab platform.
-
Single-Channel Audio Source Separation with NMF: Divergences, Constraints and Algorithms
Audio Source Separation, 2018Co-Authors: Cédric Févotte, Emmanuel Vincent, Alexey OzerovAbstract:Spectral decomposition by nonnegative matrix factorisation (NMF) has become state-of-the-art practice in many Audio signal processing tasks, such as Source Separation, enhancement or transcription. This chapter reviews the fundamentals of NMF-based Audio decomposition, in unsupervised and informed settings. We formulate NMF as an optimisation problem and discuss the choice of the measure of fit. We present the standard majorisation-minimisation strategy to address optimisation for NMF with the common $$\beta $$β-divergence, a family of measures of fit that takes the quadratic cost, the generalised Kullback-Leibler divergence and the Itakura-Saito divergence as special cases. We discuss the reconstruction of time-domain components from the spectral factorisation and present common variants of NMF-based spectral decomposition: supervised and informed settings, regularised versions, temporal models.
-
Multichannel Audio Source Separation with deep neural networks
IEEE ACM Transactions on Audio Speech and Language Processing, 2016Co-Authors: Aditya Arie Nugraha, Antoine Liutkus, Emmanuel VincentAbstract:This article addresses the problem of multichannel Audio Source Separation. We propose a framework where deep neural networks (DNNs) are used to model the Source spectra and combined with the classical multichannel Gaussian model to exploit the spatial information. The parameters are estimated in an iterative expectation-maximization (EM) fashion and used to derive a multichannel Wiener filter. We present an extensive experimental study to show the impact of different design choices on the performance of the proposed technique. We consider different cost functions for the training of DNNs, namely the probabilistically motivated Itakura-Saito divergence, and also Kullback-Leibler, Cauchy, mean squared error, and phase-sensitive cost functions. We also study the number of EM iterations and the use of multiple DNNs, where each DNN aims to improve the spectra estimated by the preceding EM iteration. Finally, we present its application to a speech enhancement problem. The experimental results show the benefit of the proposed multichannel approach over a single-channel DNN-based approach and the conventional multichannel nonnegative matrix factorization based iterative EM algorithm.
-
Fusion Methods for Speech Enhancement and Audio Source Separation
IEEE ACM Transactions on Audio Speech and Language Processing, 2016Co-Authors: Xabier Jaureguiberry, Emmanuel Vincent, Gael RichardAbstract:A wide variety of Audio Source Separation techniques exist and can already tackle many challenging industrial issues. However, in contrast with other application domains, fusion principles were rarely investigated in Audio Source Separation despite their demonstrated potential in classification tasks. In this paper, we propose a general fusion framework which takes advantage of the diversity of existing Separation techniques in order to improve Separation quality. We obtain new Source estimates by summing the individual estimates given by different Separation techniques weighted by a set of fusion coefficients. We investigate three alternative fusion methods which are based on standard nonlinear optimization, Bayesian model averaging, or deep neural networks. Experiments conducted for both speech enhancement and singing voice extraction demonstrate that all the proposed methods outperform traditional model selection. The use of deep neural networks for the estimation of time-varying coefficients notably leads to large quality improvements, up to 3 dB in terms of signal-to-distortion ratio compared to model selection.
-
multi channel Audio Source Separation using multiple deformed references
IEEE Transactions on Audio Speech and Language Processing, 2015Co-Authors: Nathan Souviraalabastie, Emmanuel Vincent, Anaik Olivero, Frédéric BimbotAbstract:We present a general multi-channel Source Separation framework where additional Audio references are available for one (or more) Source(s) of a given mixture. Each Audio reference is another mixture which is supposed to contain at least one Source similar to one of the target Sources. Deformations between the Sources of interest and their references are modeled in a linear manner using a generic formulation. This is done by adding transformation matrices to an excitation-filter model, hence affecting different axes, namely frequency, dictionary component or time. A nonnegative matrix co-factorization algorithm and a generalized expectation-maximization algorithm are used to estimate the parameters of the model. Different model parameterizations and different combinations of algorithms are tested on music plus voice mixtures guided by music and/or voice references and on professionally-produced music recordings guided by cover references. Our algorithms improve the signal-to-distortion ratio (SDR) of the Sources with the lowest intensity by 9 to 15 decibels (dB) with respect to original mixtures.
-
The Flexible Audio Source Separation Toolbox Version 2.0
2014Co-Authors: Yann Salaün, Emmanuel Vincent, Nancy Bertin, Nathan Souviraà-labastie, Xabier Jaureguiberry, Dung T. Tran, Frédéric BimbotAbstract:The Flexible Audio Source Separation Toolbox (FASST) is a toolbox for Audio Source Separation that relies on a general modeling and estimation framework that is applicable to a wide range of scenarios. We introduce the new version of the toolbox written in C++, which provides a number of advantages compared to the first Matlab version: portability, faster computation, simplified user interface, more scripting languages. In addition, we provide a state-of-the-art example of use for the Separation of speech and domestic noise. The demonstration will give attendees the opportunity to explore the settings and to experience their effect on the Separation performance.
Mark D. Plumbley - One of the best experts on this subject based on the ideXlab platform.
-
raw multi channel Audio Source Separation using multi resolution convolutional auto encoders
European Signal Processing Conference, 2018Co-Authors: Emad M. Grais, Dominic Ward, Mark D. PlumbleyAbstract:Supervised multi-channel Audio Source Separation requires extracting useful spectral, temporal, and spatial features from the mixed signals. the success of many existing systems is therefore largely dependent on the choice of features used for training. In this work, we introduce a novel multi-channel, multiresolution convolutional auto-encoder neural network that works on raw time-domain signals to determine appropriate multiresolution features for separating the singing-voice from stereo music. Our experimental results show that the proposed method can achieve multi-channel Audio Source Separation without the need for hand-crafted features or any pre- or post-processing.
-
multi resolution fully convolutional neural networks for monaural Audio Source Separation
International Conference on Latent Variable Analysis and Signal Separation, 2018Co-Authors: Emad M. Grais, Dominic Ward, Hagen Wierstorf, Mark D. PlumbleyAbstract:In deep neural networks with convolutional layers, all the neurons in each layer typically have the same size receptive fields (RFs) with the same resolution. Convolutional layers with neurons that have large RF capture global information from the input features, while layers with neurons that have small RF size capture local details with high resolution from the input features. In this work, we introduce novel deep multi-resolution fully convolutional neural networks (MR-FCN), where each layer has a range of neurons with different RF sizes to extract multi-resolution features that capture the global and local information from its input features. The proposed MR-FCN is applied to separate the singing voice from mixtures of music Sources. Experimental results show that using MR-FCN improves the performance compared to feedforward deep neural networks (DNNs) and single resolution deep fully convolutional neural networks (FCNs) on the Audio Source Separation problem.
-
multi resolution fully convolutional neural networks for monaural Audio Source Separation
arXiv: Sound, 2017Co-Authors: Emad M. Grais, Dominic Ward, Hagen Wierstorf, Mark D. PlumbleyAbstract:In deep neural networks with convolutional layers, each layer typically has fixed-size/single-resolution receptive field (RF). Convolutional layers with a large RF capture global information from the input features, while layers with small RF size capture local details with high resolution from the input features. In this work, we introduce novel deep multi-resolution fully convolutional neural networks (MR-FCNN), where each layer has different RF sizes to extract multi-resolution features that capture the global and local details information from its input features. The proposed MR-FCNN is applied to separate a target Audio Source from a mixture of many Audio Sources. Experimental results show that using MR-FCNN improves the performance compared to feedforward deep neural networks (DNNs) and single resolution deep fully convolutional neural networks (FCNNs) on the Audio Source Separation problem.
-
Estimating the loudness balance of musical mixtures using Audio Source Separation
2017Co-Authors: Dominic Ward, Mark D. Plumbley, Hagen Wierstorf, Russell Mason, Christopher HummersoneAbstract:To assist with the development of intelligent mixing systems, it would be useful to be able to extract the loudness balance of Sources in an existing musical mixture. The relative-to-mix loudness level of four instrument groups was predicted using the Sources extracted by 12 Audio Source Separation algorithms. The predictions were compared with the ground truth loudness data of the original unmixed stems obtained from a recent dataset involving 100 mixed songs. It was found that the best Source Separation system could predict the relative loudness of each instrument group with an average root-mean-square error of 1.2 LU, with superior performance obtained on vocals.
-
two stage single channel Audio Source Separation using deep neural networks
IEEE Transactions on Audio Speech and Language Processing, 2017Co-Authors: Emad M. Grais, Gerard Roma, Andrew J R Simpson, Mark D. PlumbleyAbstract:Most single channel Audio Source Separation approaches produce separated Sources accompanied by interference from other Sources and other distortions. To tackle this problem, we propose to separate the Sources in two stages. In the first stage, the Sources are separated from the mixed signal. In the second stage, the interference between the separated Sources and the distortions are reduced using deep neural networks (DNNs). We propose two methods that use DNNs to improve the quality of the separated Sources in the second stage. In the first method, each separated Source is improved individually using its own trained DNN, while in the second method all the separated Sources are improved together using a single DNN. To further improve the quality of the separated Sources, the DNNs in the second stage are trained discriminatively to further decrease the interference and the distortions of the separated Sources. Our experimental results show that using two stages of Separation improves the quality of the separated signals by decreasing the interference between the separated Sources and distortions compared to separating the Sources using a single stage of Separation.
Roland Badeau - One of the best experts on this subject based on the ideXlab platform.
-
Model-based STFT phase recovery for Audio Source Separation
IEEE Transactions on Audio Speech and Language Processing, 2018Co-Authors: Paul Magron, Roland Badeau, Bertrand DavidAbstract:For Audio Source Separation applications, it is common to estimate the magnitude of the short-time Fourier transform (STFT) of each Source. In order to further synthesizing time-domain signals, it is necessary to recover the phase of the corresponding complex-valued STFT. Most authors in this field choose a Wiener-like filtering approach which boils down to using the phase of the original mixture. In this paper, a different standpoint is adopted. Many music events are partially composed of slowly varying sinusoids and the STFT phase increment over time of those frequency components takes a specific form. This allows phase recovery by an unwrapping technique once a short-term frequency estimate has been obtained. Herein, a novel iterative Source Separation procedure is proposed which builds upon these results. It consists in minimizing the mixing error by means of the auxiliary function method. This procedure is initialized by exploiting the unwrapping technique in order to generate estimates that benefit from a temporal continuity property. Experiments conducted on realistic music pieces show that, given accurate magnitude estimates, this procedure outperforms the state-of-the-art consistent Wiener filter.
-
Model-Based STFT Phase Recovery for Audio Source Separation
IEEE Transactions on Audio Speech and Language Processing, 2018Co-Authors: Paul Magron, Roland Badeau, Bertrand DavidAbstract:For Audio Source Separation applications, it is common to estimate the magnitude of the short-time Fourier transform (STFT) of each Source. In order to further synthesize time-domain signals, it is necessary to recover the phase of the corresponding complex-valued STFT. Most authors in this field choose a Wiener-like filtering approach, which boils down to use the phase of the original mixture. In this paper, a different standpoint is adopted. Many music events are partially composed of slowly varying sinusoids and the STFT phase increment over time of those frequency components takes a specific form. This allows phase recovery by an unwrapping technique once a short-term frequency estimate has been obtained. Herein, a novel iterative Source Separation procedure is proposed that builds upon these results. It consists in minimizing the mixing error by means of the auxiliary function method. This procedure is initialized by exploiting the unwrapping technique in order to generate estimates that benefit from a temporal continuity property. Experiments conducted on realistic music pieces show that, given accurate magnitude estimates, this procedure outperforms the state-of-the-art consistent Wiener filter.
-
Phase-dependent anisotropic Gaussian model for Audio Source Separation
2017Co-Authors: Paul Magron, Roland Badeau, Bertrand DavidAbstract:Phase reconstruction of complex components in the time-frequency domain is a challenging but necessary task for Audio Source Separation. While traditional approaches do not exploit phase constraints that originate from signal modeling, some prior information about the phase can be obtained from sinusoidal modeling. In this paper, we introduce a probabilistic mixture model which allows us to incorporate such phase priors within a Source Separation framework. While the magnitudes are estimated beforehand, the phases are modeled by Von Mises random variables whose location parameters are the phase priors. We then approximate this non-tractable model by an anisotropic Gaussian model, in which the phase dependencies are preserved. This enables us to derive an MMSE estimator of the Sources which optimally combines Wiener filtering and prior phase estimates. Experimental results highlight the potential of incorporating phase priors into mixture models for separating overlapping components in complex Audio mixtures.
-
multichannel Audio Source Separation variational inference of time frequency Sources from time domain observations
International Conference on Acoustics Speech and Signal Processing, 2017Co-Authors: Simon Leglaive, Roland Badeau, Gael RichardAbstract:A great number of methods for multichannel Audio Source Separation are based on probabilistic approaches in which the Sources are modeled as latent random variables in a Time-Frequency (TF) domain. For reverberant mixtures, it is common to approximate the time-domain convolutive mixing process as being instantaneous in the short-term Fourier transform domain, under a short mixing filters assumption. The TF latent Sources are then inferred from the TF mixture observations. In this paper we propose to infer the TF latent Sources from the time-domain observations. This approach allows us to exactly model the convolutive mixing process. The inference procedure relies on a variational expectation-maximization algorithm. In significant reverberation conditions, our approach leads to a signal-to-distortion ratio improvement of 5.5 dB compared with the usual TF approximation of the convolutive mixing process.
-
informed Audio Source Separation a comparative study
European Signal Processing Conference, 2012Co-Authors: Antoine Liutkus, Roland Badeau, Laurent Girin, Laurent Daudet, Stanislaw Gorlow, Sylvain Marchand, Nicolas Sturmel, Shuhua Zhang, Gael RichardAbstract:The goal of Source Separation algorithms is to recover the constituent Sources, or Audio objects, from their mixture. However, blind algorithms still do not yield estimates of sufficient quality for many practical uses. Informed Source Separation (ISS) is a solution to make Separation robust when the Audio objects are known during a so-called encoding stage. During that stage, a small amount of side information is computed and transmitted with the mixture. At a decoding stage, when the Sources are no longer available, the mixture is processed using the side information to recover the Audio objects, thus greatly improving the quality of the estimates at a cost of additional bitrate which depends on the size of the side information. In this study, we compare six methods from the state of the art in terms of quality versus bitrate, and show that a good Separation performance can be attained at competitive bitrates.
Bertrand David - One of the best experts on this subject based on the ideXlab platform.
-
Model-based STFT phase recovery for Audio Source Separation
IEEE Transactions on Audio Speech and Language Processing, 2018Co-Authors: Paul Magron, Roland Badeau, Bertrand DavidAbstract:For Audio Source Separation applications, it is common to estimate the magnitude of the short-time Fourier transform (STFT) of each Source. In order to further synthesizing time-domain signals, it is necessary to recover the phase of the corresponding complex-valued STFT. Most authors in this field choose a Wiener-like filtering approach which boils down to using the phase of the original mixture. In this paper, a different standpoint is adopted. Many music events are partially composed of slowly varying sinusoids and the STFT phase increment over time of those frequency components takes a specific form. This allows phase recovery by an unwrapping technique once a short-term frequency estimate has been obtained. Herein, a novel iterative Source Separation procedure is proposed which builds upon these results. It consists in minimizing the mixing error by means of the auxiliary function method. This procedure is initialized by exploiting the unwrapping technique in order to generate estimates that benefit from a temporal continuity property. Experiments conducted on realistic music pieces show that, given accurate magnitude estimates, this procedure outperforms the state-of-the-art consistent Wiener filter.
-
Model-Based STFT Phase Recovery for Audio Source Separation
IEEE Transactions on Audio Speech and Language Processing, 2018Co-Authors: Paul Magron, Roland Badeau, Bertrand DavidAbstract:For Audio Source Separation applications, it is common to estimate the magnitude of the short-time Fourier transform (STFT) of each Source. In order to further synthesize time-domain signals, it is necessary to recover the phase of the corresponding complex-valued STFT. Most authors in this field choose a Wiener-like filtering approach, which boils down to use the phase of the original mixture. In this paper, a different standpoint is adopted. Many music events are partially composed of slowly varying sinusoids and the STFT phase increment over time of those frequency components takes a specific form. This allows phase recovery by an unwrapping technique once a short-term frequency estimate has been obtained. Herein, a novel iterative Source Separation procedure is proposed that builds upon these results. It consists in minimizing the mixing error by means of the auxiliary function method. This procedure is initialized by exploiting the unwrapping technique in order to generate estimates that benefit from a temporal continuity property. Experiments conducted on realistic music pieces show that, given accurate magnitude estimates, this procedure outperforms the state-of-the-art consistent Wiener filter.
-
Phase-dependent anisotropic Gaussian model for Audio Source Separation
2017Co-Authors: Paul Magron, Roland Badeau, Bertrand DavidAbstract:Phase reconstruction of complex components in the time-frequency domain is a challenging but necessary task for Audio Source Separation. While traditional approaches do not exploit phase constraints that originate from signal modeling, some prior information about the phase can be obtained from sinusoidal modeling. In this paper, we introduce a probabilistic mixture model which allows us to incorporate such phase priors within a Source Separation framework. While the magnitudes are estimated beforehand, the phases are modeled by Von Mises random variables whose location parameters are the phase priors. We then approximate this non-tractable model by an anisotropic Gaussian model, in which the phase dependencies are preserved. This enables us to derive an MMSE estimator of the Sources which optimally combines Wiener filtering and prior phase estimates. Experimental results highlight the potential of incorporating phase priors into mixture models for separating overlapping components in complex Audio mixtures.