The Experts below are selected from a list of 159543 Experts worldwide ranked by ideXlab platform

Hynek Hermansky - One of the best experts on this subject based on the ideXlab platform.

  • temporal envelope compensation for robust phoneme recognition using Modulation spectrum
    Journal of the Acoustical Society of America, 2010
    Co-Authors: Sriram Ganapathy, Samuel Thomas, Hynek Hermansky
    Abstract:

    A robust feature extraction technique for phoneme recognition is proposed which is based on deriving Modulation Frequency components from the speech signal. The Modulation Frequency components are computed from syllable-length segments of sub-band temporal envelopes estimated using Frequency domain linear prediction. Although the baseline features provide good performance in clean conditions, the performance degrades significantly in noisy conditions. In this paper, a technique for noise compensation is proposed where an estimate of the noise envelope is subtracted from the noisy speech envelope. The noise compensation technique suppresses the effect of additive noise in speech. The robustness of the proposed features is further enhanced by the gain normalization technique. The normalized temporal envelopes are compressed with static (logarithmic) and dynamic (adaptive loops) compression and are converted into Modulation Frequency features. These features are used in an automatic phoneme recognition task....

  • phoneme recognition using spectral envelope and Modulation Frequency features
    International Conference on Acoustics Speech and Signal Processing, 2009
    Co-Authors: Samuel Thomas, Sriram Ganapathy, Hynek Hermansky
    Abstract:

    We present a new feature extraction technique for phoneme recognition that uses short-term spectral envelope and Modulation Frequency features. These features are derived from sub-band temporal envelopes of speech estimated using Frequency Domain Linear Prediction (FDLP). While spectral envelope features are obtained by the short-term integration of the sub-band envelopes, the Modulation Frequency components are derived from the long-term evolution of the sub-band envelopes. These features are combined at the phoneme posterior level and used as features for a hybrid HMM-ANN phoneme recognizer. For the phoneme recognition task on the TIMIT database, the proposed features show an improvement of 4.7% over the other feature extraction techniques.

  • tandem representations of spectral envelope and Modulation Frequency features for asr
    Conference of the International Speech Communication Association, 2009
    Co-Authors: Samuel Thomas, Sriram Ganapathy, Hynek Hermansky
    Abstract:

    We present a feature extraction technique for automatic speech recognition that uses Tandem representation of short-term spectral envelope and Modulation Frequency features. These features, derived from sub-band temporal envelopes of speech estimated using Frequency domain linear prediction, are combined at the phoneme posterior level. Tandem representations derived from these phoneme posteriors are used along with HMM based ASR systems for both small and large vocabulary continuous speech recognition (LVCSR) tasks. For a small vocabulary continuous digit task on the OGI Digits database, the proposed features reduce the word error rate (WER) by 13 % relative to other feature extraction techniques. We obtain a relative reduction of about 14 % in WER for an LVCSR task using the NIST RT05 evaluation data. For phoneme recognition tasks on the TIMIT database these features provide a relative improvement of 13% compared to other techniques.

  • Modulation Frequency features for phoneme recognition in noisy speech
    Journal of the Acoustical Society of America, 2009
    Co-Authors: Sriram Ganapathy, Samuel Thomas, Hynek Hermansky
    Abstract:

    In this letter, a new feature extraction technique based on Modulation spectrum derived from syllable-length segments of subband temporal envelopes is proposed. These subband envelopes are derived from autoregressive modeling of Hilbert envelopes of the signal in critical bands, processed by both a static (logarithmic) and a dynamic (adaptive loops) compression. These features are then used for machine recognition of phonemes in telephone speech. Without degrading the performance in clean conditions, the proposed features show significant improvements compared to other state-of-the-art speech analysis techniques. In addition to the overall phoneme recognition rates, the performance with broad phonetic classes is reported.

C S Pattichis - One of the best experts on this subject based on the ideXlab platform.

  • despeckle filtering for multiscale amplitude Modulation Frequency Modulation am fm texture analysis of ultrasound images of the intima media complex
    International Journal of Biomedical Imaging, 2014
    Co-Authors: Christos P Loizou, Victor Murray, Marios S Pattichis, M Pantziaris, A N Nicolaides, C S Pattichis
    Abstract:

    The intima-media thickness (IMT) of the common carotid artery (CCA) is widely used as an early indicator of cardiovascular disease (CVD). Typically, the IMT grows with age and this is used as a sign of increased risk of CVD. Beyond thickness, there is also clinical interest in identifying how the composition and texture of the intima-media complex (IMC) changed and how these textural changes grow into atherosclerotic plaques that can cause stroke. Clearly though texture analysis of ultrasound images can be greatly affected by speckle noise, our goal here is to develop effective despeckle noise methods that can recover image texture associated with increased rates of atherosclerosis disease. In this study, we perform a comparative evaluation of several despeckle filtering methods, on 100 ultrasound images of the CCA, based on the extracted multiscale Amplitude-Modulation Frequency-Modulation (AM-FM) texture features and visual image quality assessment by two clinical experts. Texture features were extracted from the automatically segmented IMC for three different age groups. The despeckle filters hybrid median and the homogeneous mask area filter showed the best performance by improving the class separation between the three age groups and also yielded significantly improved image quality.

  • multiscale amplitude Modulation Frequency Modulation am fm texture analysis of ultrasound images of the intima and media layers of the carotid artery
    International Conference of the IEEE Engineering in Medicine and Biology Society, 2011
    Co-Authors: Christos P Loizou, Victor Murray, Marios S Pattichis, M Pantziaris, C S Pattichis
    Abstract:

    The intima-media thickness (IMT) of the common carotid artery (CCA) is widely used as an early indicator of cardiovascular disease (CVD). Clinically, there is strong interest in identifying how the composition and texture of the media layer (ML) can be associated with the risk of stroke. In this study, we use 2-D amplitude-Modulation Frequency-Modulation (AM-FM) analysis of the intima-media complex (IMC), the ML, and intima layer (IL) of the CCA to detect texture changes as a function of age and sex. The study was performed on 100 ultrasound images acquired from asymptomatic subjects at risk of atherosclerosis. To investigate texture variations associated with age, we separated them into three age groups: 1) patients younger than 50; 2) patients aged between 50 and 60 years old; and 3) patients over 60 years old. We also separated the patients by sex. The IMC, ML, and IL were segmented manually by a neurovascular expert and also by a snake-based segmentation system. To reject strong edge artifacts, we prefilter with an AM-FM filterbank that is centered along the horizontal Frequency axis (parallel to the long axis of the IMC, ML, and IL), while removing the low-pass filter estimates and Frequency bands with large, vertical Frequency components. To investigate significant texture changes, we extract the instantaneous amplitude (IA) and the magnitude of the instantaneous Frequency (IF) over each layer component, for low-, medium-, and high-Frequency AM-FM components. We detected significant texture differences between the higher risk age group of >;60 years versus the lower risk age group of ;60 groups, we found significant differences in the medium-scale IA extracted from the IMC. Between the >;60 and the 50-60 groups, we found significant texture changes in the low scale IA and high-scale IF magnitude extracted from the IMC, and the low-scale IA extracted from the IL. Also, we noted that the IA for the ML showed significant differences between males and females for all age groups. The AM-FM features provide complimentary information to classical texture analysis features like the gray-scale median, contrast, and coarseness. These findings provide evidence that AM-FM texture features can be associated with the progression of cardiovascular risk for disease and the risk of stroke with age. However, a larger scale study is needed to establish the application in clinical practice.

  • multiscale amplitude Modulation Frequency Modulation am fm texture analysis of multiple sclerosis in brain mri images
    International Conference of the IEEE Engineering in Medicine and Biology Society, 2011
    Co-Authors: Christos P Loizou, Victor Murray, Marios S Pattichis, M Pantziaris, Ioannis Seimenis, C S Pattichis
    Abstract:

    This study introduces the use of multiscale amplitude Modulation-Frequency Modulation (AM-FM) texture analysis of multiple sclerosis (MS) using magnetic resonance (MR) images from brain. Clinically, there is interest in identifying potential associations between lesion texture and disease progression, and in relating texture features with relevant clinical indexes, such as the expanded disability status scale (EDSS). This longitudinal study explores the application of 2-D AM-FM analysis of brain white matter MS lesions to quantify and monitor disease load. To this end, MS lesions and normal-appearing white matter (NAWM) from MS patients, as well as normal white matter (NWM) from healthy volunteers, were segmented on transverse T2-weighted images obtained from serial brain MR imaging (MRI) scans (0 and 6-12 months). The instantaneous amplitude (IA), the magnitude of the instantaneous Frequency (IF), and the IF angle were extracted from each segmented region at different scales. The findings suggest that AM-FM characteristics succeed in differentiating 1) between NWM and lesions; 2) between NAWM and lesions; and 3) between NWM and NAWM. A support vector machine (SVM) classifier succeeded in differentiating between patients that, two years after the initial MRI scan, acquired an EDSS ≤ 2 from those with EDSS >; 2 (correct classification rate = 86%). The best classification results were obtained from including the combination of the low-scale IA and IF magnitude with the medium-scale IA. The AM-FM features provide complementary information to classical texture analysis features like the gray-scale median, contrast, and coarseness. The findings of this study provide evidence that AM-FM features may have a potential role as surrogate markers of lesion load in MS.

Qianjie Fu - One of the best experts on this subject based on the ideXlab platform.

  • envelope interactions in multi channel amplitude Modulation Frequency discrimination by cochlear implant users
    PLOS ONE, 2015
    Co-Authors: Deniz Baskent, John J. Galvin, Monita Chatterjee, Qianjie Fu
    Abstract:

    Rationale Previous cochlear implant (CI) studies have shown that single-channel amplitude Modulation Frequency discrimination (AMFD) can be improved when coherent Modulation is delivered to additional channels. It is unclear whether the multi-channel advantage is due to increased loudness, multiple envelope representations, or to component channels with better temporal processing. Measuring envelope interference may shed light on how modulated channels can be combined. Methods In this study, multi-channel AMFD was measured in CI subjects using a 3-alternative forced-choice, non-adaptive procedure ("which interval is different?"). For the reference stimulus, the reference AM (100 Hz) was delivered to all 3 channels. For the probe stimulus, the target AM (101, 102, 104, 108, 116, 132, 164, 228, or 256 Hz) was delivered to 1 of 3 channels, and the reference AM (100 Hz) delivered to the other 2 channels. The spacing between electrodes was varied to be wide or narrow to test different degrees of channel interaction. Results Results showed that CI subjects were highly sensitive to interactions between the reference and target envelopes. However, performance was non-monotonic as a function of target AM Frequency. For the wide spacing, there was significantly less envelope interaction when the target AM was delivered to the basal channel. For the narrow spacing, there was no effect of target AM channel. The present data were also compared to a related previous study in which the target AM was delivered to a single channel or to all 3 channels. AMFD was much better with multiple than with single channels whether the target AM was delivered to 1 of 3 or to all 3 channels. For very small differences between the reference and target AM frequencies (2-4 Hz), there was often greater sensitivity when the target AM was delivered to 1 of 3 channels versus all 3 channels, especially for narrowly spaced electrodes. Conclusions Besides the increased loudness, the present results also suggest that multiple envelope representations may contribute to the multi-channel advantage observed in previous AMFD studies. The different patterns of results for the wide and narrow spacing suggest a peripheral contribution to multi-channel temporal processing. Because the effect of target AM Frequency was non-monotonic in this study, adaptive procedures may not be suitable to measure AMFD thresholds with interfering envelopes. Envelope interactions among multiple channels may be quite complex, depending on the envelope information presented to each channel and the relative independence of the stimulated channels.

  • Modulation Frequency discrimination with single and multiple channels in cochlear implant users
    Hearing Research, 2015
    Co-Authors: John J. Galvin, Deniz Baskent, Qianjie Fu
    Abstract:

    Temporal envelope cues convey important speech information for cochlear implant (CI) users. Many studies have explored CI users' single-channel temporal envelope processing. However, in clinical CI speech processors, temporal envelope information is processed by multiple channels. Previous studies have shown that amplitude Modulation Frequency discrimination (AMFD) thresholds are better when temporal envelopes are delivered to multiple rather than single channels. In clinical fitting, current levels on single channels must often be reduced to accommodate multi-channel loudness summation. As such, it is unclear whether the multi-channel advantage in AMFD observed in previous studies was due to coherent envelope information distributed across the cochlea or to greater loudness associated with multi-channel stimulation. In this study, single- and multi-channel AMFD thresholds were measured in CI users. Multi-channel component electrodes were either widely or narrowly spaced to vary the degree of overlap between neural populations. The reference amplitude Modulation (AM) Frequency was 100 Hz, and coherent Modulation was applied to all channels. In Experiment 1, single- and multi-channel AMFD thresholds were measured at similar loudness. In this case, current levels on component channels were higher for single-than for multi-channel AM stimuli, and the Modulation depth was approximately 100% of the perceptual dynamic range (i.e., between threshold and maximum acceptable loudness). Results showed no significant difference in AMFD thresholds between similarly loud single- and multi-channel modulated stimuli. In Experiment 2, single- and multi-channel AMFD thresholds were compared at substantially different loudness. In this case, current levels on component channels were the same for single- and multi-channel stimuli (“summation-adjusted” current levels) and the same range of Modulation (in dB) was applied to the component channels for both single- and multi-channel testing. With the summation-adjusted current levels, loudness was lower with single than with multiple channels and the AM depth resulted in substantial stimulation below single-channel audibility, thereby reducing the perceptual range of AM. Results showed that AMFD thresholds were significantly better with multiple channels than with any of the single component channels. There was no significant effect of the distribution of electrodes on multi-channel AMFD thresholds. The results suggest that increased loudness due to multi-channel summation may contribute to the multi-channel advantage in AMFD, and that overall loudness may matter more than the distribution of envelope information in the cochlea.

John J. Galvin - One of the best experts on this subject based on the ideXlab platform.

  • envelope interactions in multi channel amplitude Modulation Frequency discrimination by cochlear implant users
    PLOS ONE, 2015
    Co-Authors: Deniz Baskent, John J. Galvin, Monita Chatterjee, Qianjie Fu
    Abstract:

    Rationale Previous cochlear implant (CI) studies have shown that single-channel amplitude Modulation Frequency discrimination (AMFD) can be improved when coherent Modulation is delivered to additional channels. It is unclear whether the multi-channel advantage is due to increased loudness, multiple envelope representations, or to component channels with better temporal processing. Measuring envelope interference may shed light on how modulated channels can be combined. Methods In this study, multi-channel AMFD was measured in CI subjects using a 3-alternative forced-choice, non-adaptive procedure ("which interval is different?"). For the reference stimulus, the reference AM (100 Hz) was delivered to all 3 channels. For the probe stimulus, the target AM (101, 102, 104, 108, 116, 132, 164, 228, or 256 Hz) was delivered to 1 of 3 channels, and the reference AM (100 Hz) delivered to the other 2 channels. The spacing between electrodes was varied to be wide or narrow to test different degrees of channel interaction. Results Results showed that CI subjects were highly sensitive to interactions between the reference and target envelopes. However, performance was non-monotonic as a function of target AM Frequency. For the wide spacing, there was significantly less envelope interaction when the target AM was delivered to the basal channel. For the narrow spacing, there was no effect of target AM channel. The present data were also compared to a related previous study in which the target AM was delivered to a single channel or to all 3 channels. AMFD was much better with multiple than with single channels whether the target AM was delivered to 1 of 3 or to all 3 channels. For very small differences between the reference and target AM frequencies (2-4 Hz), there was often greater sensitivity when the target AM was delivered to 1 of 3 channels versus all 3 channels, especially for narrowly spaced electrodes. Conclusions Besides the increased loudness, the present results also suggest that multiple envelope representations may contribute to the multi-channel advantage observed in previous AMFD studies. The different patterns of results for the wide and narrow spacing suggest a peripheral contribution to multi-channel temporal processing. Because the effect of target AM Frequency was non-monotonic in this study, adaptive procedures may not be suitable to measure AMFD thresholds with interfering envelopes. Envelope interactions among multiple channels may be quite complex, depending on the envelope information presented to each channel and the relative independence of the stimulated channels.

  • Modulation Frequency discrimination with single and multiple channels in cochlear implant users
    Hearing Research, 2015
    Co-Authors: John J. Galvin, Deniz Baskent, Qianjie Fu
    Abstract:

    Temporal envelope cues convey important speech information for cochlear implant (CI) users. Many studies have explored CI users' single-channel temporal envelope processing. However, in clinical CI speech processors, temporal envelope information is processed by multiple channels. Previous studies have shown that amplitude Modulation Frequency discrimination (AMFD) thresholds are better when temporal envelopes are delivered to multiple rather than single channels. In clinical fitting, current levels on single channels must often be reduced to accommodate multi-channel loudness summation. As such, it is unclear whether the multi-channel advantage in AMFD observed in previous studies was due to coherent envelope information distributed across the cochlea or to greater loudness associated with multi-channel stimulation. In this study, single- and multi-channel AMFD thresholds were measured in CI users. Multi-channel component electrodes were either widely or narrowly spaced to vary the degree of overlap between neural populations. The reference amplitude Modulation (AM) Frequency was 100 Hz, and coherent Modulation was applied to all channels. In Experiment 1, single- and multi-channel AMFD thresholds were measured at similar loudness. In this case, current levels on component channels were higher for single-than for multi-channel AM stimuli, and the Modulation depth was approximately 100% of the perceptual dynamic range (i.e., between threshold and maximum acceptable loudness). Results showed no significant difference in AMFD thresholds between similarly loud single- and multi-channel modulated stimuli. In Experiment 2, single- and multi-channel AMFD thresholds were compared at substantially different loudness. In this case, current levels on component channels were the same for single- and multi-channel stimuli (“summation-adjusted” current levels) and the same range of Modulation (in dB) was applied to the component channels for both single- and multi-channel testing. With the summation-adjusted current levels, loudness was lower with single than with multiple channels and the AM depth resulted in substantial stimulation below single-channel audibility, thereby reducing the perceptual range of AM. Results showed that AMFD thresholds were significantly better with multiple channels than with any of the single component channels. There was no significant effect of the distribution of electrodes on multi-channel AMFD thresholds. The results suggest that increased loudness due to multi-channel summation may contribute to the multi-channel advantage in AMFD, and that overall loudness may matter more than the distribution of envelope information in the cochlea.

Samuel Thomas - One of the best experts on this subject based on the ideXlab platform.

  • temporal envelope compensation for robust phoneme recognition using Modulation spectrum
    Journal of the Acoustical Society of America, 2010
    Co-Authors: Sriram Ganapathy, Samuel Thomas, Hynek Hermansky
    Abstract:

    A robust feature extraction technique for phoneme recognition is proposed which is based on deriving Modulation Frequency components from the speech signal. The Modulation Frequency components are computed from syllable-length segments of sub-band temporal envelopes estimated using Frequency domain linear prediction. Although the baseline features provide good performance in clean conditions, the performance degrades significantly in noisy conditions. In this paper, a technique for noise compensation is proposed where an estimate of the noise envelope is subtracted from the noisy speech envelope. The noise compensation technique suppresses the effect of additive noise in speech. The robustness of the proposed features is further enhanced by the gain normalization technique. The normalized temporal envelopes are compressed with static (logarithmic) and dynamic (adaptive loops) compression and are converted into Modulation Frequency features. These features are used in an automatic phoneme recognition task....

  • phoneme recognition using spectral envelope and Modulation Frequency features
    International Conference on Acoustics Speech and Signal Processing, 2009
    Co-Authors: Samuel Thomas, Sriram Ganapathy, Hynek Hermansky
    Abstract:

    We present a new feature extraction technique for phoneme recognition that uses short-term spectral envelope and Modulation Frequency features. These features are derived from sub-band temporal envelopes of speech estimated using Frequency Domain Linear Prediction (FDLP). While spectral envelope features are obtained by the short-term integration of the sub-band envelopes, the Modulation Frequency components are derived from the long-term evolution of the sub-band envelopes. These features are combined at the phoneme posterior level and used as features for a hybrid HMM-ANN phoneme recognizer. For the phoneme recognition task on the TIMIT database, the proposed features show an improvement of 4.7% over the other feature extraction techniques.

  • tandem representations of spectral envelope and Modulation Frequency features for asr
    Conference of the International Speech Communication Association, 2009
    Co-Authors: Samuel Thomas, Sriram Ganapathy, Hynek Hermansky
    Abstract:

    We present a feature extraction technique for automatic speech recognition that uses Tandem representation of short-term spectral envelope and Modulation Frequency features. These features, derived from sub-band temporal envelopes of speech estimated using Frequency domain linear prediction, are combined at the phoneme posterior level. Tandem representations derived from these phoneme posteriors are used along with HMM based ASR systems for both small and large vocabulary continuous speech recognition (LVCSR) tasks. For a small vocabulary continuous digit task on the OGI Digits database, the proposed features reduce the word error rate (WER) by 13 % relative to other feature extraction techniques. We obtain a relative reduction of about 14 % in WER for an LVCSR task using the NIST RT05 evaluation data. For phoneme recognition tasks on the TIMIT database these features provide a relative improvement of 13% compared to other techniques.

  • Modulation Frequency features for phoneme recognition in noisy speech
    Journal of the Acoustical Society of America, 2009
    Co-Authors: Sriram Ganapathy, Samuel Thomas, Hynek Hermansky
    Abstract:

    In this letter, a new feature extraction technique based on Modulation spectrum derived from syllable-length segments of subband temporal envelopes is proposed. These subband envelopes are derived from autoregressive modeling of Hilbert envelopes of the signal in critical bands, processed by both a static (logarithmic) and a dynamic (adaptive loops) compression. These features are then used for machine recognition of phonemes in telephone speech. Without degrading the performance in clean conditions, the proposed features show significant improvements compared to other state-of-the-art speech analysis techniques. In addition to the overall phoneme recognition rates, the performance with broad phonetic classes is reported.