The Experts below are selected from a list of 360 Experts worldwide ranked by ideXlab platform

Févotte Cédric - One of the best experts on this subject based on the ideXlab platform.

  • Phase retrieval with Bregman divergences and application to Audio Signal recovery
    'Institute of Electrical and Electronics Engineers (IEEE)', 2021
    Co-Authors: Vial Pierre-hugo, Magron Paul, Oberlin Thomas, Févotte Cédric
    Abstract:

    Phase retrieval (PR) aims to recover a Signal from the magnitudes of a set of inner products. This problem arises in many Audio Signal Processing applications which operate on a short-time Fourier transform magnitude or power spectrogram, and discard the phase information. Recovering the missing phase from the resulting modified spectrogram is indeed necessary in order to synthesize time-domain Signals. PR is commonly addressed by considering a minimization problem involving a quadratic loss function. In this paper, we adopt a different standpoint. Indeed, the quadratic loss does not properly account for some perceptual properties of Audio, and alternative discrepancy measures such as beta-divergences have been preferred in many settings. Therefore, we formulate PR as a new minimization problem involving Bregman divergences. Since these divergences are not symmetric with respect to their two input arguments in general, they lead to two different formulations of the problem. To optimize the resulting objective, we derive two algorithms based on accelerated gradient descent and alternating direction method of multipliers. Experiments conducted on Audio Signal recovery from spectrograms that are either exact or estimated from noisy observations highlight the potential of our proposed methods for Audio restoration. In particular, leveraging some of these Bregman divergences induce better performance than the quadratic loss when performing PR from spectrograms under very noisy conditions.Comment: 23 pages, 3 figures, accepted for publication in the IEEE Journal of Selected Topics in Signal Processin

  • Phase retrieval with Bregman divergences and application to Audio Signal recovery
    2020
    Co-Authors: Vial Pierre-hugo, Magron Paul, Oberlin Thomas, Févotte Cédric
    Abstract:

    Phase retrieval (PR) aims to recover a Signal from the magnitudes of a set of inner products. This problem arises in many Audio Signal Processing applications which operate on a short-time Fourier transform magnitude or power spectrogram, and discard the phase information. Recovering the missing phase from the resulting modified spectrogram is indeed necessary in order to synthesize time-domain Signals. PR is commonly addressed by considering a minimization problem involving a quadratic loss function. In this paper, we adopt a different standpoint. Indeed, the quadratic loss does not properly account for some perceptual properties of Audio, and alternative discrepancy measures such as beta-divergences have been preferred in many settings. Therefore, we formulate PR as a new minimization problem involving Bregman divergences. We consider a general formulation that actually addresses two problems, since it accounts for the non-symmetry of these divergences in general. To optimize the resulting objective, we derive two algorithms based on accelerated gradient descent and alternating direction method of multiplier. Experiments conducted on Audio Signal recovery from either exact or modified spectrograms highlight the potential of our proposed methods for Audio restoration. In particular, leveraging some of these Bregman divergences induce better performance than the quadratic loss when performing PR from highly degraded spectrograms.Comment: 23 pages, 4 figures, submitted to the IEEE Journal of Selected Topics in Signal Processin

  • Phase retrieval with Bregman divergences: Application to Audio Signal recovery
    2020
    Co-Authors: Vial Pierre-hugo, Magron Paul, Oberlin Thomas, Févotte Cédric
    Abstract:

    Phase retrieval aims to recover a Signal from magnitude or power spectra measurements. It is often addressed by considering a minimization problem involving a quadratic cost function. We propose a different formulation based on Bregman divergences, which encompass divergences that are appropriate for Audio Signal Processing applications. We derive a fast gradient algorithm to solve this problem.Comment: in Proceedings of iTWIST'20, Paper-ID: 16, Nantes, France, December, 2-4, 202

Alexey Ozerov - One of the best experts on this subject based on the ideXlab platform.

  • Single-Channel Audio Source Separation with NMF: Divergences, Constraints and Algorithms
    Audio Source Separation, 2018
    Co-Authors: Cédric Févotte, Emmanuel Vincent, Alexey Ozerov
    Abstract:

    Spectral decomposition by nonnegative matrix factorisation (NMF) has become state-of-the-art practice in many Audio Signal Processing tasks, such as source separation, enhancement or transcription. This chapter reviews the fundamentals of NMF-based Audio decomposition, in unsupervised and informed settings. We formulate NMF as an optimisation problem and discuss the choice of the measure of fit. We present the standard majorisation-minimisation strategy to address optimisation for NMF with the common $$\beta $$β-divergence, a family of measures of fit that takes the quadratic cost, the generalised Kullback-Leibler divergence and the Itakura-Saito divergence as special cases. We discuss the reconstruction of time-domain components from the spectral factorisation and present common variants of NMF-based spectral decomposition: supervised and informed settings, regularised versions, temporal models.

  • a consolidated perspective on multimicrophone speech enhancement and source separation
    IEEE Transactions on Audio Speech and Language Processing, 2017
    Co-Authors: Sharon Gannot, Emmanuel Vincent, Shmulik Markovichgolan, Alexey Ozerov
    Abstract:

    Speech enhancement and separation are core problems in Audio Signal Processing, with commercial applications in devices as diverse as mobile phones, conference call systems, hands-free systems, or hearing aids. In addition, they are crucial preProcessing steps for noise-robust automatic speech and speaker recognition. Many devices now have two to eight microphones. The enhancement and separation capabilities offered by these multichannel interfaces are usually greater than those of single-channel interfaces. Research in speech enhancement and separation has followed two convergent paths, starting with microphone array Processing and blind source separation, respectively. These communities are now strongly interrelated and routinely borrow ideas from each other. Yet, a comprehensive overview of the common foundations and the differences between these approaches is lacking at present. In this paper, we propose to fill this gap by analyzing a large number of established and recent techniques according to four transverse axes: 1 the acoustic impulse response model, 2 the spatial filter design criterion, 3 the parameter estimation algorithm, and 4 optional postfiltering. We conclude this overview paper by providing a list of software and data resources and by discussing perspectives and future trends in the field.

Tara N. Sainath - One of the best experts on this subject based on the ideXlab platform.

  • Deep Learning for Audio Signal Processing
    IEEE Journal of Selected Topics in Signal Processing, 2019
    Co-Authors: Hendrik Purwins, Shuo-yiin Chang, Jan Schlüter, Tuomas Virtanen, Bo Li, Tara N. Sainath
    Abstract:

    Given the recent surge in developments of deep learning, this paper provides a review of the state-of-the-art deep learning techniques for Audio Signal Processing. Speech, music, and environmental sound Processing are considered side-by-side, in order to point out similarities and differences between the domains, highlighting general methods, problems, key references, and potential for cross fertilization between areas. The dominant feature representations (in particular, log-mel spectra and raw waveform) and deep learning models are reviewed, including convolutional neural networks, variants of the long short-term memory architecture, as well as more Audio-specific neural network models. Subsequently, prominent deep learning application areas are covered, i.e., Audio recognition (automatic speech recognition, music information retrieval, environmental sound detection, localization and tracking) and synthesis and transformation (source separation, Audio enhancement, generative models for speech, sound, and music synthesis). Finally, key issues and future questions regarding deep learning applied to Audio Signal Processing are identified.

James D Johnston - One of the best experts on this subject based on the ideXlab platform.

  • Nonuniform oversampled filter banks for Audio Signal Processing
    IEEE Transactions on Speech and Audio Processing, 2003
    Co-Authors: Zoran Cvetković, James D Johnston
    Abstract:

    In emerging Audio technology applications, there is a need for decompositions of Audio Signals into oversampled subband components with time-frequency resolution which mimics that of the cochlear filter bank and with high aliasing attenuation in each of the subbands independently, rather than aliasing cancellation properties. We present a design of nearly perfect reconstruction nonuniform oversampled filter banks which implement Signal decompositions of this kind.

Sharon Gannot - One of the best experts on this subject based on the ideXlab platform.

  • a consolidated perspective on multimicrophone speech enhancement and source separation
    IEEE Transactions on Audio Speech and Language Processing, 2017
    Co-Authors: Sharon Gannot, Emmanuel Vincent, Shmulik Markovichgolan, Alexey Ozerov
    Abstract:

    Speech enhancement and separation are core problems in Audio Signal Processing, with commercial applications in devices as diverse as mobile phones, conference call systems, hands-free systems, or hearing aids. In addition, they are crucial preProcessing steps for noise-robust automatic speech and speaker recognition. Many devices now have two to eight microphones. The enhancement and separation capabilities offered by these multichannel interfaces are usually greater than those of single-channel interfaces. Research in speech enhancement and separation has followed two convergent paths, starting with microphone array Processing and blind source separation, respectively. These communities are now strongly interrelated and routinely borrow ideas from each other. Yet, a comprehensive overview of the common foundations and the differences between these approaches is lacking at present. In this paper, we propose to fill this gap by analyzing a large number of established and recent techniques according to four transverse axes: 1 the acoustic impulse response model, 2 the spatial filter design criterion, 3 the parameter estimation algorithm, and 4 optional postfiltering. We conclude this overview paper by providing a list of software and data resources and by discussing perspectives and future trends in the field.

  • spatial source subtraction based on incomplete measurements of relative transfer function
    IEEE Transactions on Audio Speech and Language Processing, 2015
    Co-Authors: Zbyněk Koldovský, Jiři Malek, Sharon Gannot
    Abstract:

    Relative impulse responses between microphones are usually long and dense due to the reverberant acoustic environment. Estimating them from short and noisy recordings poses a long-standing challenge of Audio Signal Processing. In this paper, we apply a novel strategy based on ideas of compressed sensing. Relative transfer function (RTF) corresponding to the relative impulse response can often be estimated accurately from noisy data but only for certain frequencies. This means that often only an incomplete measurement of the RTF is available. A complete RTF estimate can be obtained through finding its sparsest representation in the time-domain: that is, through computing the sparsest among the corresponding relative impulse responses. Based on this approach, we propose to estimate the RTF from noisy data in three steps. First, the RTF is estimated using any conventional method such as the nonstationarity-based estimator by Gannot et al. or through blind source separation. Second, frequencies are determined for which the RTF estimate appears to be accurate. Third, the RTF is reconstructed through solving a weighted l 1 convex program, which we propose to solve via a computationally efficient variant of the SpaRSA (Sparse Reconstruction by Separable Approximation) algorithm. An extensive experimental study with real-world recordings has been conducted. It has been shown that the proposed method is capable of improving many conventional estimators used as the first step in most situations.