The Experts below are selected from a list of 10020 Experts worldwide ranked by ideXlab platform

Yoichi Haneda - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP - Diffused sensing for sharp directivity microphone array
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Okamoto Manabu, Yoichi Haneda
    Abstract:

    We propose a method for achieving sharp directivity by sensing signals in a diffuse acoustic field. Directivity control based on a beamforming method has been studied to make it possible to extract the waveform and location of an identified target source even if there are many noise sources. Sharp directivity can be achieved by minimizing the output noise power of a beamforming filter. However, it is difficult to minimize the output noise power over a broad frequency ranges. Our approach for minimizing the output noise power is to control the Spatial properties of the transfer functions and the Spatial Correlation Matrix, by using a reflector that surrounds a microphone array. We investigated the relationships between the output noise power and the structure of the Spatial Correlation Matrix and found that it was possible to minimize the output noise power by sensing diffuse acoustic signals and by designing filters taking the diffuseness of the acoustic field into consideration. In experiments, we observed diffusely reflected signals by placing a truncated-octahedral reflector near a spherical microphone array. We designed filters by using measured transfer functions and confirmed that the proposed method was effective for reducing the output noise power and forming a sharp directivity beamforming filter.

  • estimating direct to reverberant energy ratio using d r Spatial Correlation Matrix model
    IEEE Transactions on Audio Speech and Language Processing, 2011
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound Spatial Correlation Matrix model (Hereafter referred to as the Spatial Correlation model). This model expresses the Spatial Correlation Matrix of an array input signal as two Spatial Correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the Spatial Correlation Matrix of the measured signal using the Spatial Correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method.

  • estimating direct to reverberant energy ratio based on Spatial Correlation model segregating direct sound and reverberation
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    A new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on amodel of a Spatial Correlation Matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the Correlation Matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved.

  • estimation of sound source orientation using eigenspace of Spatial Correlation Matrix
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important Spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of Spatial Correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.

  • ICASSP - Estimation of sound source orientation using eigenspace of Spatial Correlation Matrix
    2010 IEEE International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important Spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of Spatial Correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.

Kenta Niwa - One of the best experts on this subject based on the ideXlab platform.

  • Optimal Microphone Array Observation for Clear Recording of Distant Sound Sources
    IEEE ACM Transactions on Audio Speech and Language Processing, 2016
    Co-Authors: Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi
    Abstract:

    We propose the principle for deriving an optimum design for a microphone array that uses mutual information to segregate distant sound sources. Many conventional studies on array signal processing have focused on methods for estimating sound sources from array observations. To record distant sound sources clearly, designing an optimum array structure to segregate a target from other noise is also necessary. In this study, we reveal that the optimum array observation was achieved by receiving signals that are physically decorrelated between microphones, which homogenizes the eigenvalues of the Spatial Correlation Matrix. We theoretically explain this underlying principle using mutual information between sound sources and microphone observations whose relation to the existing minimum mean square error criterion for source separation is also discussed. The implementation of such a microphone array is possible by placing microphones in front of parabolic reflectors since the phase/amplitude around the focal point of the reflectors drastically varies with small perturbation of the microphone position. CrossCorrelation between observed signals can be reduced by optimally placing microphones. An array structure based on our proposed principle was tested by implementing minimum variance distortion-less response beamforming and postfiltering in the observations of a prototype microphone array. We experimentally confirmed that 1) the eigenvalues of the Spatial Correlation Matrix were asymptotically homogenized and 2) the target source could be extracted clearly even when the sound sources were positioned 16.5 m from the array.

  • ICASSP - Microphone array for increasing mutual information between sound sources and observation signals
    2015 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2015
    Co-Authors: Kenta Niwa, Tatsuya Kako, Kobayashi Kazunori
    Abstract:

    We investigated the basic principle of how Spatial signals should be captured with a microphone array to estimate each source signal and its practical implementation. Most conventional studies on array signal processing have been focused on the design of beamforming and Wiener filters. To achieve further effective noise reduction, designing an optimum array structure to segregate a target from other noises is necessary. We found the optimum structure of the Spatial Correlation Matrix to estimate each source signal. This is achieved by receiving signals whose eigenvalues of the Spatial Correlation Matrix are homogenized. To homogenize the eigenvalues of the Spatial Correlation Matrix while maintaining a short impulse response length, we propose an array structure composed of parabolic reflectors and 96 microphones. Through experiments using the proposed array structure, we confirmed that the eigenvalues of the Spatial Correlation Matrix was asymptotically homogenized and that sharp directivity could be formed.

  • ICASSP - Diffused sensing for sharp directivity microphone array
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Okamoto Manabu, Yoichi Haneda
    Abstract:

    We propose a method for achieving sharp directivity by sensing signals in a diffuse acoustic field. Directivity control based on a beamforming method has been studied to make it possible to extract the waveform and location of an identified target source even if there are many noise sources. Sharp directivity can be achieved by minimizing the output noise power of a beamforming filter. However, it is difficult to minimize the output noise power over a broad frequency ranges. Our approach for minimizing the output noise power is to control the Spatial properties of the transfer functions and the Spatial Correlation Matrix, by using a reflector that surrounds a microphone array. We investigated the relationships between the output noise power and the structure of the Spatial Correlation Matrix and found that it was possible to minimize the output noise power by sensing diffuse acoustic signals and by designing filters taking the diffuseness of the acoustic field into consideration. In experiments, we observed diffusely reflected signals by placing a truncated-octahedral reflector near a spherical microphone array. We designed filters by using measured transfer functions and confirmed that the proposed method was effective for reducing the output noise power and forming a sharp directivity beamforming filter.

  • estimating direct to reverberant energy ratio using d r Spatial Correlation Matrix model
    IEEE Transactions on Audio Speech and Language Processing, 2011
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound Spatial Correlation Matrix model (Hereafter referred to as the Spatial Correlation model). This model expresses the Spatial Correlation Matrix of an array input signal as two Spatial Correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the Spatial Correlation Matrix of the measured signal using the Spatial Correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method.

  • estimating direct to reverberant energy ratio based on Spatial Correlation model segregating direct sound and reverberation
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    A new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on amodel of a Spatial Correlation Matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the Correlation Matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved.

Kenichi Furuya - One of the best experts on this subject based on the ideXlab platform.

  • CISIS - Reducing Computational Complexity of Multichannel Nonnegative Matrix Factorization Using Initial Value Setting for Speech Recognition
    Advances in Intelligent Systems and Computing, 2018
    Co-Authors: Taiki Izumi, Ryo Aihara, Yohei Okato, Takanobu Uramoto, Shingo Uenohara, Toshiyuki Hanazawa, Kenichi Furuya
    Abstract:

    In this paper, we propose efficient the number of computational iteration method of MNMF for speech recognition. The proposed method initializes estimates MNMF algorithm with the estimated Spatial Correlation Matrix reduces the number of iteration of updates algorithm. The experiment result shows that our method reduced the computational complexity of MNMF.

  • APSIPA - Multichannel NMF with Reduced Computational Complexity for Speech Recognition
    2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2018
    Co-Authors: Taiki Izumi, Ryo Aihara, Takanobu Uramoto, Shingo Uenohara, Toshiyuki Hanazawa, Kenichi Furuya, Yohei Okato
    Abstract:

    In this study, we propose efficient the number of computational iteration method of MNMF for speech recognition. The proposed method initializes and estimates the MNMF algorithm with respect to the estimated Spatial Correlation Matrix reducing the number of iteration of update algorithm. This time, mask emphasis via Expectation Maximization algorithm is used for estimation of a Spatial Correlation Matrix. As another method, we propose a computational complexity reduction method via decimating update of the Spatial Correlation MatrixH. The experimental result indicates that our method reduced the computational complexity of MNMF. It shows that the performance of the conventional MNMF was maintained and the computational complexity could be reduced.

  • ICASSP - Diffused sensing for sharp directivity microphone array
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Okamoto Manabu, Yoichi Haneda
    Abstract:

    We propose a method for achieving sharp directivity by sensing signals in a diffuse acoustic field. Directivity control based on a beamforming method has been studied to make it possible to extract the waveform and location of an identified target source even if there are many noise sources. Sharp directivity can be achieved by minimizing the output noise power of a beamforming filter. However, it is difficult to minimize the output noise power over a broad frequency ranges. Our approach for minimizing the output noise power is to control the Spatial properties of the transfer functions and the Spatial Correlation Matrix, by using a reflector that surrounds a microphone array. We investigated the relationships between the output noise power and the structure of the Spatial Correlation Matrix and found that it was possible to minimize the output noise power by sensing diffuse acoustic signals and by designing filters taking the diffuseness of the acoustic field into consideration. In experiments, we observed diffusely reflected signals by placing a truncated-octahedral reflector near a spherical microphone array. We designed filters by using measured transfer functions and confirmed that the proposed method was effective for reducing the output noise power and forming a sharp directivity beamforming filter.

  • estimating direct to reverberant energy ratio using d r Spatial Correlation Matrix model
    IEEE Transactions on Audio Speech and Language Processing, 2011
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound Spatial Correlation Matrix model (Hereafter referred to as the Spatial Correlation model). This model expresses the Spatial Correlation Matrix of an array input signal as two Spatial Correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the Spatial Correlation Matrix of the measured signal using the Spatial Correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method.

  • estimating direct to reverberant energy ratio based on Spatial Correlation model segregating direct sound and reverberation
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    A new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on amodel of a Spatial Correlation Matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the Correlation Matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved.

Yusuke Hioka - One of the best experts on this subject based on the ideXlab platform.

  • Optimal Microphone Array Observation for Clear Recording of Distant Sound Sources
    IEEE ACM Transactions on Audio Speech and Language Processing, 2016
    Co-Authors: Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi
    Abstract:

    We propose the principle for deriving an optimum design for a microphone array that uses mutual information to segregate distant sound sources. Many conventional studies on array signal processing have focused on methods for estimating sound sources from array observations. To record distant sound sources clearly, designing an optimum array structure to segregate a target from other noise is also necessary. In this study, we reveal that the optimum array observation was achieved by receiving signals that are physically decorrelated between microphones, which homogenizes the eigenvalues of the Spatial Correlation Matrix. We theoretically explain this underlying principle using mutual information between sound sources and microphone observations whose relation to the existing minimum mean square error criterion for source separation is also discussed. The implementation of such a microphone array is possible by placing microphones in front of parabolic reflectors since the phase/amplitude around the focal point of the reflectors drastically varies with small perturbation of the microphone position. CrossCorrelation between observed signals can be reduced by optimally placing microphones. An array structure based on our proposed principle was tested by implementing minimum variance distortion-less response beamforming and postfiltering in the observations of a prototype microphone array. We experimentally confirmed that 1) the eigenvalues of the Spatial Correlation Matrix were asymptotically homogenized and 2) the target source could be extracted clearly even when the sound sources were positioned 16.5 m from the array.

  • estimating direct to reverberant energy ratio using d r Spatial Correlation Matrix model
    IEEE Transactions on Audio Speech and Language Processing, 2011
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound Spatial Correlation Matrix model (Hereafter referred to as the Spatial Correlation model). This model expresses the Spatial Correlation Matrix of an array input signal as two Spatial Correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the Spatial Correlation Matrix of the measured signal using the Spatial Correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method.

  • estimating direct to reverberant energy ratio based on Spatial Correlation model segregating direct sound and reverberation
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    A new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on amodel of a Spatial Correlation Matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the Correlation Matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved.

  • estimation of sound source orientation using eigenspace of Spatial Correlation Matrix
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important Spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of Spatial Correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.

  • ICASSP - Estimation of sound source orientation using eigenspace of Spatial Correlation Matrix
    2010 IEEE International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important Spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of Spatial Correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.

Sumitaka Sakauchi - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP - Diffused sensing for sharp directivity microphone array
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Okamoto Manabu, Yoichi Haneda
    Abstract:

    We propose a method for achieving sharp directivity by sensing signals in a diffuse acoustic field. Directivity control based on a beamforming method has been studied to make it possible to extract the waveform and location of an identified target source even if there are many noise sources. Sharp directivity can be achieved by minimizing the output noise power of a beamforming filter. However, it is difficult to minimize the output noise power over a broad frequency ranges. Our approach for minimizing the output noise power is to control the Spatial properties of the transfer functions and the Spatial Correlation Matrix, by using a reflector that surrounds a microphone array. We investigated the relationships between the output noise power and the structure of the Spatial Correlation Matrix and found that it was possible to minimize the output noise power by sensing diffuse acoustic signals and by designing filters taking the diffuseness of the acoustic field into consideration. In experiments, we observed diffusely reflected signals by placing a truncated-octahedral reflector near a spherical microphone array. We designed filters by using measured transfer functions and confirmed that the proposed method was effective for reducing the output noise power and forming a sharp directivity beamforming filter.

  • estimating direct to reverberant energy ratio using d r Spatial Correlation Matrix model
    IEEE Transactions on Audio Speech and Language Processing, 2011
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound Spatial Correlation Matrix model (Hereafter referred to as the Spatial Correlation model). This model expresses the Spatial Correlation Matrix of an array input signal as two Spatial Correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the Spatial Correlation Matrix of the measured signal using the Spatial Correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method.

  • estimating direct to reverberant energy ratio based on Spatial Correlation model segregating direct sound and reverberation
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    A new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on amodel of a Spatial Correlation Matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the Correlation Matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved.

  • estimation of sound source orientation using eigenspace of Spatial Correlation Matrix
    International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important Spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of Spatial Correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.

  • ICASSP - Estimation of sound source orientation using eigenspace of Spatial Correlation Matrix
    2010 IEEE International Conference on Acoustics Speech and Signal Processing, 2010
    Co-Authors: Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Kenichi Furuya, Yoichi Haneda
    Abstract:

    We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important Spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of Spatial Correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.