The Experts below are selected from a list of 120654 Experts worldwide ranked by ideXlab platform

Herve Bourlard - One of the best experts on this subject based on the ideXlab platform.

  • robust hmm based speech music segmentation
    International Conference on Acoustics Speech and Signal Processing, 2002
    Co-Authors: Jitendra Ajmera, Iain Mccowan, Herve Bourlard
    Abstract:

    In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the Local Probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.

  • ICASSP - Robust HMM-based speech/music segmentation
    IEEE International Conference on Acoustics Speech and Signal Processing, 2002
    Co-Authors: Jitendra Ajmera, Iain A. Mccowan, Herve Bourlard
    Abstract:

    In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the Local Probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.

Jitendra Ajmera - One of the best experts on this subject based on the ideXlab platform.

  • robust hmm based speech music segmentation
    International Conference on Acoustics Speech and Signal Processing, 2002
    Co-Authors: Jitendra Ajmera, Iain Mccowan, Herve Bourlard
    Abstract:

    In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the Local Probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.

  • ICASSP - Robust HMM-based speech/music segmentation
    IEEE International Conference on Acoustics Speech and Signal Processing, 2002
    Co-Authors: Jitendra Ajmera, Iain A. Mccowan, Herve Bourlard
    Abstract:

    In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the Local Probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.

S Gazor - One of the best experts on this subject based on the ideXlab platform.

  • Local Probability distribution of natural signals in sparse domains
    International Journal of Adaptive Control and Signal Processing, 2013
    Co-Authors: Hossein Rabbani, S Gazor
    Abstract:

    SUMMARY In this paper, we investigate the Local PDF of natural signals in sparse domains. The statistical properties of natural signals are characterized more accurately in the sparse domains because the sparse domain coefficients have heavy-tailed distribution and have reduced correlation with adjacent coefficients. Our experiments on 3D data in 3D discrete complex wavelet transform domain show that a conditionally (given Locally estimated variance and shape) independent Bessel K-form distribution (BKFD) Locally fits the sparse domain's coefficients of natural signals, accurately. To justify this observation, we also investigate the PDF of the Locally estimated variance and suggest a Gamma PDF for the Locally estimated variance. Because commonly used sparse transformations are orthonormal, the PDF of the sparse domain coefficients must converge to Gaussian distribution by virtue of central limit theorem assuming that natural signals are Locally wide sense stationary for small window sizes. Interestingly, we observe that the PDF of the normalized data (on the Locally estimated variance) exhibit a Gaussian PDF, which confirms that the BKFD is an appropriate fit. Copyright © 2013 John Wiley & Sons, Ltd.

  • Local Probability distribution of natural signals in sparse domains
    International Conference on Acoustics Speech and Signal Processing, 2011
    Co-Authors: Hossein Rabbani, S Gazor
    Abstract:

    In this paper we investigate the Local Probability density function (pdf) of natural signals in sparse domains. The statistical properties of natural signals are characterized more accurately in the sparse domains because the sparse domain coefficients (SDCs) have heavy-tailed distribution and have reduced correlation with adjacent coefficients. Our experiments show that a conditionally (given Locally estimated variance and shape) independent Bessel K-form (BKF) pdf Locally fits the sparse domain's coefficients of natural signals, accurately. To justify this observation, we also investigate the pdf of the Locally estimated variance and suggest a Gamma pdf for the Locally estimated variance. Since commonly used sparse transformations are orthonormal, the pdf of the sparse domain coefficients must converge to Gaussian distribution by virtue of central limit theorem assuming that natural signals are Locally wide sense stationary for small window sizes. Interestingly, we observe that the pdf of the normalized data (on the Locally estimated variance) exhibit a Gaussian pdf, which justifies why the BKF pdf is an appropriate fit.

  • ICASSP - Local Probability distribution of natural signals in sparse domains
    2011 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2011
    Co-Authors: Hossein Rabbani, S Gazor
    Abstract:

    In this paper we investigate the Local Probability density function (pdf) of natural signals in sparse domains. The statistical properties of natural signals are characterized more accurately in the sparse domains because the sparse domain coefficients (SDCs) have heavy-tailed distribution and have reduced correlation with adjacent coefficients. Our experiments show that a conditionally (given Locally estimated variance and shape) independent Bessel K-form (BKF) pdf Locally fits the sparse domain's coefficients of natural signals, accurately. To justify this observation, we also investigate the pdf of the Locally estimated variance and suggest a Gamma pdf for the Locally estimated variance. Since commonly used sparse transformations are orthonormal, the pdf of the sparse domain coefficients must converge to Gaussian distribution by virtue of central limit theorem assuming that natural signals are Locally wide sense stationary for small window sizes. Interestingly, we observe that the pdf of the normalized data (on the Locally estimated variance) exhibit a Gaussian pdf, which justifies why the BKF pdf is an appropriate fit.

Hossein Rabbani - One of the best experts on this subject based on the ideXlab platform.

  • Local Probability distribution of natural signals in sparse domains
    International Journal of Adaptive Control and Signal Processing, 2013
    Co-Authors: Hossein Rabbani, S Gazor
    Abstract:

    SUMMARY In this paper, we investigate the Local PDF of natural signals in sparse domains. The statistical properties of natural signals are characterized more accurately in the sparse domains because the sparse domain coefficients have heavy-tailed distribution and have reduced correlation with adjacent coefficients. Our experiments on 3D data in 3D discrete complex wavelet transform domain show that a conditionally (given Locally estimated variance and shape) independent Bessel K-form distribution (BKFD) Locally fits the sparse domain's coefficients of natural signals, accurately. To justify this observation, we also investigate the PDF of the Locally estimated variance and suggest a Gamma PDF for the Locally estimated variance. Because commonly used sparse transformations are orthonormal, the PDF of the sparse domain coefficients must converge to Gaussian distribution by virtue of central limit theorem assuming that natural signals are Locally wide sense stationary for small window sizes. Interestingly, we observe that the PDF of the normalized data (on the Locally estimated variance) exhibit a Gaussian PDF, which confirms that the BKFD is an appropriate fit. Copyright © 2013 John Wiley & Sons, Ltd.

  • Local Probability distribution of natural signals in sparse domains
    International Conference on Acoustics Speech and Signal Processing, 2011
    Co-Authors: Hossein Rabbani, S Gazor
    Abstract:

    In this paper we investigate the Local Probability density function (pdf) of natural signals in sparse domains. The statistical properties of natural signals are characterized more accurately in the sparse domains because the sparse domain coefficients (SDCs) have heavy-tailed distribution and have reduced correlation with adjacent coefficients. Our experiments show that a conditionally (given Locally estimated variance and shape) independent Bessel K-form (BKF) pdf Locally fits the sparse domain's coefficients of natural signals, accurately. To justify this observation, we also investigate the pdf of the Locally estimated variance and suggest a Gamma pdf for the Locally estimated variance. Since commonly used sparse transformations are orthonormal, the pdf of the sparse domain coefficients must converge to Gaussian distribution by virtue of central limit theorem assuming that natural signals are Locally wide sense stationary for small window sizes. Interestingly, we observe that the pdf of the normalized data (on the Locally estimated variance) exhibit a Gaussian pdf, which justifies why the BKF pdf is an appropriate fit.

  • ICASSP - Local Probability distribution of natural signals in sparse domains
    2011 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2011
    Co-Authors: Hossein Rabbani, S Gazor
    Abstract:

    In this paper we investigate the Local Probability density function (pdf) of natural signals in sparse domains. The statistical properties of natural signals are characterized more accurately in the sparse domains because the sparse domain coefficients (SDCs) have heavy-tailed distribution and have reduced correlation with adjacent coefficients. Our experiments show that a conditionally (given Locally estimated variance and shape) independent Bessel K-form (BKF) pdf Locally fits the sparse domain's coefficients of natural signals, accurately. To justify this observation, we also investigate the pdf of the Locally estimated variance and suggest a Gamma pdf for the Locally estimated variance. Since commonly used sparse transformations are orthonormal, the pdf of the sparse domain coefficients must converge to Gaussian distribution by virtue of central limit theorem assuming that natural signals are Locally wide sense stationary for small window sizes. Interestingly, we observe that the pdf of the normalized data (on the Locally estimated variance) exhibit a Gaussian pdf, which justifies why the BKF pdf is an appropriate fit.

Iain Mccowan - One of the best experts on this subject based on the ideXlab platform.

  • robust hmm based speech music segmentation
    International Conference on Acoustics Speech and Signal Processing, 2002
    Co-Authors: Jitendra Ajmera, Iain Mccowan, Herve Bourlard
    Abstract:

    In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the Local Probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.