The Experts below are selected from a list of 2604 Experts worldwide ranked by ideXlab platform

Hiroshi Saruwatari - One of the best experts on this subject based on the ideXlab platform.

  • dnn based frequency component prediction for frequency domain Audio source separation
    European Signal Processing Conference, 2021
    Co-Authors: Rui Watanabe, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
    Abstract:

    Multichannel Audio source separation (MASS) plays an important role in various Audio applications. Frequency-domain MASS algorithms such as Multichannel nonnegative matrix factorization achieve better separation quality. However, they require a considerable computational cost for estimating the frequency-wise separation filter. To solve this problem, we propose a new framework combining the MASS algorithms and a simple deep neural network (DNN). In the proposed framework, frequency-domain MASS is performed only in narrowband frequency bins. Then, DNN predicts the separated source components in other frequency bins, where both the observed mixture of all frequency bins and the separated narrowband source components are used as DNN inputs. Our experimental results show the validity of the proposed MASS framework in terms of computational efficiency.

  • independent deeply learned matrix analysis with automatic selection of stable microphone wise update and fast sourcewise update of demixing matrix
    Signal Processing, 2021
    Co-Authors: Naoki Makishima, Daichi Kitamura, Norihiro Takamune, Hiroshi Saruwatari, Yu Takahashi, Yoshiki Mitsui, Kazunobu Kondo
    Abstract:

    Abstract Independent deeply learned matrix analysis (IDLMA) is a fast and high-performance method for Multichannel Audio source separation. IDLMA utilizes the deep neural network inference of source models and the blind estimation of demixing filters based on source independence. In conventional IDLMA, iterative projection (IP) is exploited to estimate the demixing filters. Although IP is a fast algorithm, it sometimes fails to estimate an appropriate solution. This is because IP updates the demixing filters in a sourcewise manner, where only one source model is used for each update, and the update sometimes becomes unstable owing to the specific low-quality source models. In this paper, we first derive a new numerically stable microphone-wise update algorithm that exploits all source model information simultaneously. The microphone-wise update problem cannot be solved by IP; instead, a new type of vectorwise coordinate descent algorithm is introduced. Next, comparison analysis of the proposed microphone-wise update and IP reveals the tradeoff w.r.t. convergence speed and numerical stability. To resolve this tradeoff problem, we propose the automatic selection of update rules on the basis of the likelihood function of observed signals. Finally, experimental results show the efficacy of the proposed IDLMA with the automatic selection of update rules.

  • independent deeply learned matrix analysis for determined Audio source separation
    IEEE Transactions on Audio Speech and Language Processing, 2019
    Co-Authors: Naoki Makishima, Shinichi Mogami, Hayato Sumino, Daichi Kitamura, Norihiro Takamune, Shinnosuke Takamichi, Hiroshi Saruwatari
    Abstract:

    In this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based Multichannel Audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded Audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy.

  • Independent Deeply Learned Matrix Analysis for Multichannel Audio Source Separation
    arXiv: Audio and Speech Processing, 2018
    Co-Authors: Shinichi Mogami, Hayato Sumino, Daichi Kitamura, Norihiro Takamune, Shinnosuke Takamichi, Hiroshi Saruwatari
    Abstract:

    In this paper, we address a Multichannel Audio source separation task and propose a new efficient method called independent deeply learned matrix analysis (IDLMA). IDLMA estimates the demixing matrix in a blind manner and updates the time-frequency structures of each source using a pretrained deep neural network (DNN). Also, we introduce a complex Student's t-distribution as a generalized source generative model including both complex Gaussian and Cauchy distributions. Experiments are conducted using music signals with a training dataset, and the results show the validity of the proposed method in terms of separation accuracy and computational cost.

  • temporal quantization of spatial information using directional clustering for Multichannel Audio coding
    Workshop on Applications of Signal Processing to Audio and Acoustics, 2009
    Co-Authors: Shigeki Miyabe, Hiroshi Saruwatari, Keisuke Masatoki, Kiyohiro Shikano, Toshiyuki Nomura
    Abstract:

    Binaural cue coding, which is a representing low bit-rate coding of Multichannel Audio, generates large distortion when the Audio data have complex spatial image, such as symphony. Such distortion caused by the low frequency resolution of spatial information because BCC quantizes the parameters of localization. In this paper we propose a new coding framework by quantizing the spatial information temporally. The single-channel sum signal is panned to the multiple channels by selecting the prototypes of the spatial filter. Optimization of the prototypes with minimum coding error is given by a k-means-like clustering of the angles whose centroids are given by the first principal components of the covariances in the classes. The efficiency of the proposed coding with high quality is verified both in the objective and subjective evaluations.

Cédric Févotte - One of the best experts on this subject based on the ideXlab platform.

  • notes on nonnegative tensor factorization of the spectrogram for Audio source separation statistical insights and towards self clustering of the spatial cues
    Computer Music Modeling and Retrieval, 2010
    Co-Authors: Cédric Févotte, Alexey Ozerov
    Abstract:

    Nonnegative tensor factorization (NTF) of Multichannel spectrograms under PARAFAC structure has recently been proposed by Fitzgerald et al as a mean of performing blind source separation (BSS) of Multichannel Audio data. In this paper we investigate the statistical source models implied by this approach. We show that it implicitly assumes a nonpoint-source model contrasting with usual BSS assumptions and we clarify the links between the measure of fit chosen for the NTF and the implied statistical distribution of the sources. While the original approach of Fitzgeral et al requires a posterior clustering of the spatial cues to group the NTF components into sources, we discuss means of performing the clustering within the factorization. In the results section we test the impact of the simplifying nonpoint-source assumption on underdetermined linear instantaneous mixtures of musical sources and discuss the limits of the approach for such mixtures.

  • Multichannel nonnegative matrix factorization in convolutive mixtures for Audio source separation
    IEEE Transactions on Audio Speech and Language Processing, 2010
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    We consider inference in a general data-driven object-based model of Multichannel Audio data, assumed generated as a possibly underdetermined convolutive mixture of source signals. We work in the short-time Fourier transform (STFT) domain, where convolution is routinely approximated as linear instantaneous mixing in each frequency band. Each source STFT is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura-Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likelihood of the Multichannel data using an expectation-maximization (EM) algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decomposition algorithms are applied to stereo Audio source separation in various settings, covering blind and supervised separation, music and speech sources, synthetic instantaneous and convolutive mixtures, as well as professionally produced music recordings. Our EM method produces competitive results with respect to state-of-the-art as illustrated on two tasks from the international Signal Separation Evaluation Campaign (SiSEC 2008).

  • C.: Multichannel nonnegative matrix factorization in convolutive mixtures for Audio source separation
    2010
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    We consider inference in a general data-driven object-based model of Multichannel Audio data, assumed generated as a possibly underdetermined convolutive mixture of source signals. Each source is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura-Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likelihood of the Multichannel data using an expectation-maximization algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decomposition algorithms were applied to stereo music and assessed in terms of blind source separation performance. Index Terms — Multichannel Audio, nonnegative matrix factorization, nonnegative tensor factorization, underdetermined convolutive blind source separation. 1

  • Multichannel nonnegative matrix factorization in convolutive mixtures with application to blind Audio source separation
    International Conference on Acoustics Speech and Signal Processing, 2009
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    We consider inference in a general data-driven object-based model of Multichannel Audio data, assumed generated as a possibly under-determined convolutive mixture of source signals. Each source is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura-Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likelihood of the Multichannel data using an expectation-maximization algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decomposition algorithms were applied to stereo music and assessed in terms of blind source separation performance.

  • Multichannel nonnegative matrix factorization in convolutive mixtures. With application to blind Audio source separation
    2009
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    Abstract—We consider inference in a general data-driven ob-ject-based model of Multichannel Audio data, assumed generated as a possibly underdetermined convolutive mixture of source signals. We work in the short-time Fourier transform (STFT) domain, where convolution is routinely approximated as linear instantaneous mixing in each frequency band. Each source STFT is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura–Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likeli-hood of the Multichannel data using an expectation-maximization (EM) algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decom-position algorithms are applied to stereo Audio source separation in various settings, covering blind and supervised separation, music and speech sources, synthetic instantaneous and convolutive mixtures, as well as professionally produced music recordings. Our EM method produces competitive results with respect to state-of-the-art as illustrated on two tasks from the international Signal Separation Evaluation Campaign (SiSEC 2008). Index Terms—Expectation-maximization (EM) algorithm, Multichannel Audio, nonnegative matrix factorization (NMF), nonnegative tensor factorization (NTF), underdetermined convo-lutive blind source separation (BSS). I

Daichi Kitamura - One of the best experts on this subject based on the ideXlab platform.

  • dnn based frequency component prediction for frequency domain Audio source separation
    European Signal Processing Conference, 2021
    Co-Authors: Rui Watanabe, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
    Abstract:

    Multichannel Audio source separation (MASS) plays an important role in various Audio applications. Frequency-domain MASS algorithms such as Multichannel nonnegative matrix factorization achieve better separation quality. However, they require a considerable computational cost for estimating the frequency-wise separation filter. To solve this problem, we propose a new framework combining the MASS algorithms and a simple deep neural network (DNN). In the proposed framework, frequency-domain MASS is performed only in narrowband frequency bins. Then, DNN predicts the separated source components in other frequency bins, where both the observed mixture of all frequency bins and the separated narrowband source components are used as DNN inputs. Our experimental results show the validity of the proposed MASS framework in terms of computational efficiency.

  • independent deeply learned matrix analysis with automatic selection of stable microphone wise update and fast sourcewise update of demixing matrix
    Signal Processing, 2021
    Co-Authors: Naoki Makishima, Daichi Kitamura, Norihiro Takamune, Hiroshi Saruwatari, Yu Takahashi, Yoshiki Mitsui, Kazunobu Kondo
    Abstract:

    Abstract Independent deeply learned matrix analysis (IDLMA) is a fast and high-performance method for Multichannel Audio source separation. IDLMA utilizes the deep neural network inference of source models and the blind estimation of demixing filters based on source independence. In conventional IDLMA, iterative projection (IP) is exploited to estimate the demixing filters. Although IP is a fast algorithm, it sometimes fails to estimate an appropriate solution. This is because IP updates the demixing filters in a sourcewise manner, where only one source model is used for each update, and the update sometimes becomes unstable owing to the specific low-quality source models. In this paper, we first derive a new numerically stable microphone-wise update algorithm that exploits all source model information simultaneously. The microphone-wise update problem cannot be solved by IP; instead, a new type of vectorwise coordinate descent algorithm is introduced. Next, comparison analysis of the proposed microphone-wise update and IP reveals the tradeoff w.r.t. convergence speed and numerical stability. To resolve this tradeoff problem, we propose the automatic selection of update rules on the basis of the likelihood function of observed signals. Finally, experimental results show the efficacy of the proposed IDLMA with the automatic selection of update rules.

  • independent deeply learned matrix analysis for determined Audio source separation
    IEEE Transactions on Audio Speech and Language Processing, 2019
    Co-Authors: Naoki Makishima, Shinichi Mogami, Hayato Sumino, Daichi Kitamura, Norihiro Takamune, Shinnosuke Takamichi, Hiroshi Saruwatari
    Abstract:

    In this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based Multichannel Audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded Audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy.

  • Independent Deeply Learned Matrix Analysis for Multichannel Audio Source Separation
    arXiv: Audio and Speech Processing, 2018
    Co-Authors: Shinichi Mogami, Hayato Sumino, Daichi Kitamura, Norihiro Takamune, Shinnosuke Takamichi, Hiroshi Saruwatari
    Abstract:

    In this paper, we address a Multichannel Audio source separation task and propose a new efficient method called independent deeply learned matrix analysis (IDLMA). IDLMA estimates the demixing matrix in a blind manner and updates the time-frequency structures of each source using a pretrained deep neural network (DNN). Also, we introduce a complex Student's t-distribution as a generalized source generative model including both complex Gaussian and Cauchy distributions. Experiments are conducted using music signals with a training dataset, and the results show the validity of the proposed method in terms of separation accuracy and computational cost.

Kazunobu Kondo - One of the best experts on this subject based on the ideXlab platform.

  • dnn based frequency component prediction for frequency domain Audio source separation
    European Signal Processing Conference, 2021
    Co-Authors: Rui Watanabe, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
    Abstract:

    Multichannel Audio source separation (MASS) plays an important role in various Audio applications. Frequency-domain MASS algorithms such as Multichannel nonnegative matrix factorization achieve better separation quality. However, they require a considerable computational cost for estimating the frequency-wise separation filter. To solve this problem, we propose a new framework combining the MASS algorithms and a simple deep neural network (DNN). In the proposed framework, frequency-domain MASS is performed only in narrowband frequency bins. Then, DNN predicts the separated source components in other frequency bins, where both the observed mixture of all frequency bins and the separated narrowband source components are used as DNN inputs. Our experimental results show the validity of the proposed MASS framework in terms of computational efficiency.

  • independent deeply learned matrix analysis with automatic selection of stable microphone wise update and fast sourcewise update of demixing matrix
    Signal Processing, 2021
    Co-Authors: Naoki Makishima, Daichi Kitamura, Norihiro Takamune, Hiroshi Saruwatari, Yu Takahashi, Yoshiki Mitsui, Kazunobu Kondo
    Abstract:

    Abstract Independent deeply learned matrix analysis (IDLMA) is a fast and high-performance method for Multichannel Audio source separation. IDLMA utilizes the deep neural network inference of source models and the blind estimation of demixing filters based on source independence. In conventional IDLMA, iterative projection (IP) is exploited to estimate the demixing filters. Although IP is a fast algorithm, it sometimes fails to estimate an appropriate solution. This is because IP updates the demixing filters in a sourcewise manner, where only one source model is used for each update, and the update sometimes becomes unstable owing to the specific low-quality source models. In this paper, we first derive a new numerically stable microphone-wise update algorithm that exploits all source model information simultaneously. The microphone-wise update problem cannot be solved by IP; instead, a new type of vectorwise coordinate descent algorithm is introduced. Next, comparison analysis of the proposed microphone-wise update and IP reveals the tradeoff w.r.t. convergence speed and numerical stability. To resolve this tradeoff problem, we propose the automatic selection of update rules on the basis of the likelihood function of observed signals. Finally, experimental results show the efficacy of the proposed IDLMA with the automatic selection of update rules.

Alexey Ozerov - One of the best experts on this subject based on the ideXlab platform.

  • notes on nonnegative tensor factorization of the spectrogram for Audio source separation statistical insights and towards self clustering of the spatial cues
    Computer Music Modeling and Retrieval, 2010
    Co-Authors: Cédric Févotte, Alexey Ozerov
    Abstract:

    Nonnegative tensor factorization (NTF) of Multichannel spectrograms under PARAFAC structure has recently been proposed by Fitzgerald et al as a mean of performing blind source separation (BSS) of Multichannel Audio data. In this paper we investigate the statistical source models implied by this approach. We show that it implicitly assumes a nonpoint-source model contrasting with usual BSS assumptions and we clarify the links between the measure of fit chosen for the NTF and the implied statistical distribution of the sources. While the original approach of Fitzgeral et al requires a posterior clustering of the spatial cues to group the NTF components into sources, we discuss means of performing the clustering within the factorization. In the results section we test the impact of the simplifying nonpoint-source assumption on underdetermined linear instantaneous mixtures of musical sources and discuss the limits of the approach for such mixtures.

  • Multichannel nonnegative matrix factorization in convolutive mixtures for Audio source separation
    IEEE Transactions on Audio Speech and Language Processing, 2010
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    We consider inference in a general data-driven object-based model of Multichannel Audio data, assumed generated as a possibly underdetermined convolutive mixture of source signals. We work in the short-time Fourier transform (STFT) domain, where convolution is routinely approximated as linear instantaneous mixing in each frequency band. Each source STFT is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura-Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likelihood of the Multichannel data using an expectation-maximization (EM) algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decomposition algorithms are applied to stereo Audio source separation in various settings, covering blind and supervised separation, music and speech sources, synthetic instantaneous and convolutive mixtures, as well as professionally produced music recordings. Our EM method produces competitive results with respect to state-of-the-art as illustrated on two tasks from the international Signal Separation Evaluation Campaign (SiSEC 2008).

  • C.: Multichannel nonnegative matrix factorization in convolutive mixtures for Audio source separation
    2010
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    We consider inference in a general data-driven object-based model of Multichannel Audio data, assumed generated as a possibly underdetermined convolutive mixture of source signals. Each source is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura-Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likelihood of the Multichannel data using an expectation-maximization algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decomposition algorithms were applied to stereo music and assessed in terms of blind source separation performance. Index Terms — Multichannel Audio, nonnegative matrix factorization, nonnegative tensor factorization, underdetermined convolutive blind source separation. 1

  • Multichannel nonnegative matrix factorization in convolutive mixtures with application to blind Audio source separation
    International Conference on Acoustics Speech and Signal Processing, 2009
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    We consider inference in a general data-driven object-based model of Multichannel Audio data, assumed generated as a possibly under-determined convolutive mixture of source signals. Each source is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura-Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likelihood of the Multichannel data using an expectation-maximization algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decomposition algorithms were applied to stereo music and assessed in terms of blind source separation performance.

  • Multichannel nonnegative matrix factorization in convolutive mixtures. With application to blind Audio source separation
    2009
    Co-Authors: Alexey Ozerov, Cédric Févotte
    Abstract:

    Abstract—We consider inference in a general data-driven ob-ject-based model of Multichannel Audio data, assumed generated as a possibly underdetermined convolutive mixture of source signals. We work in the short-time Fourier transform (STFT) domain, where convolution is routinely approximated as linear instantaneous mixing in each frequency band. Each source STFT is given a model inspired from nonnegative matrix factorization (NMF) with the Itakura–Saito divergence, which underlies a statistical model of superimposed Gaussian components. We address estimation of the mixing and source parameters using two methods. The first one consists of maximizing the exact joint likeli-hood of the Multichannel data using an expectation-maximization (EM) algorithm. The second method consists of maximizing the sum of individual likelihoods of all channels using a multiplicative update algorithm inspired from NMF methodology. Our decom-position algorithms are applied to stereo Audio source separation in various settings, covering blind and supervised separation, music and speech sources, synthetic instantaneous and convolutive mixtures, as well as professionally produced music recordings. Our EM method produces competitive results with respect to state-of-the-art as illustrated on two tasks from the international Signal Separation Evaluation Campaign (SiSEC 2008). Index Terms—Expectation-maximization (EM) algorithm, Multichannel Audio, nonnegative matrix factorization (NMF), nonnegative tensor factorization (NTF), underdetermined convo-lutive blind source separation (BSS). I