The Experts below are selected from a list of 303 Experts worldwide ranked by ideXlab platform

A Gersho - One of the best experts on this subject based on the ideXlab platform.

  • a multiple description Speech Coder based on amr wb for mobile ad hoc networks
    International Conference on Acoustics Speech and Signal Processing, 2004
    Co-Authors: Hui Dong, J D Gibson, A Gersho, V. Cuperman
    Abstract:

    To address the challenging task of achieving effective voice communication over mobile ad hoc networks (MANETs), we introduce a new multiple description (MD) Speech Coder based on the AMR-WB (adaptive multirate wideband) standard. The MD Coder splits the bitstream of the AMR-WB Coder into two redundant sub-streams by directly selecting overlapping subsets of encoded data generated for each frame. The sub-streams are transmitted on different network paths. When both sub-streams arrive at the deCoder, an output identical to that of AMR-WB is recovered. If only one substream arrives at the deCoder, degraded, but still acceptable, Speech quality is obtained. The performance of the multiple description coding system is tested for MANETS with a network simulator and an informal listening test. The results demonstrate that this approach makes effective use of the channel capacity and provides reliable end-to-end connections for voice communication in MANETs.

  • enhanced waveform interpolative coding at low bit rate
    IEEE Transactions on Speech and Audio Processing, 2001
    Co-Authors: O Gottesman, A Gersho
    Abstract:

    This paper presents a high quality enhanced waveform interpolative (EWI) Speech Coder at low bit-rate. The system incorporates novel features such as optimization of the slowly evolving waveform (SEW) for interpolation, analysis-by-synthesis (AbS) vector quantization (VQ) of the SEW dispersion phase, dual-predictive AbS quantization of the SEW, efficient parameterization of the rapidly-evolving waveform (REW) magnitude, and VQ of the REW parameter, a special pitch search for transitions, and switched-predictive analysis-by-synthesis gain VQ. Subjective tests indicate that the 2.8 kb/s EWI Coder's quality exceeds that of G.723.1 at 5.3 kb/s, and it is slightly better than that of G.723.1 at 6.3 kb/s.

  • a 1200 bps Speech Coder based on melp
    International Conference on Acoustics Speech and Signal Processing, 2000
    Co-Authors: Tian Wang, V. Cuperman, Kazuhito Koishida, A Gersho, J S Collura
    Abstract:

    This paper presents a 1.2 kbps Speech Coder based on the mixed excitation linear prediction (MELP) analysis algorithm. In the proposed Coder, the MELP parameters of three consecutive frames are grouped into a superframe and jointly quantized to obtain a high coding efficiency. The interframe redundancy is exploited with distinct quantization schemes for different unvoiced/voiced (U/V) frame combinations in the superframe. Novel techniques for improving performance make use of the superframe structure. These include pitch vector quantization using pitch differentials, joint quantization of pitch and U/V decisions and LSF quantization with a forward-backward interpolation method. Subjective test results indicate that the 1.2 kbps Speech Coder achieves approximately the same quality as the proposed federal standard 2.4 kbps MELP Coder.

  • high quality enhanced waveform interpolative coding at 2 8 kbps
    International Conference on Acoustics Speech and Signal Processing, 2000
    Co-Authors: O Gottesman, A Gersho
    Abstract:

    This paper presents a high quality enhanced waveform interpolative (EWI) Speech Coder at 2.8 kbps. The system incorporates novel features such as: dual-predictive analysis-by-synthesis (AbS) quantization of the slowly evolving waveform (SEW), efficient parametrization of the rapidly evolving waveform (REW) magnitude, and AbS vector quantization (VQ) of the REW parameter. Subjective tests indicate that its quality exceeds that of G.723.1 at 5.3 kbps, and it is slightly better than that of G.723.1 at 6.3 kbps.

  • enhanced waveform interpolative coding at 4 kbps
    1999 IEEE Workshop on Speech Coding Proceedings. Model Coders and Error Criteria (Cat. No.99EX351), 1999
    Co-Authors: O Gottesman, A Gersho
    Abstract:

    This paper presents an enhanced waveform interpolative (EWI) Speech Coder at 4 kbps. The system incorporates novel features such as analysis-by-synthesis (AbS) vector-quantization (VQ) of the dispersion-phase, AbS optimization of the slowly evolving waveform (SEW), a special pitch search for transitions, and switched-predictive analysis-by-synthesis gain VQ. Subjective quality tests indicate that it exceeds that of MPEG-4 at 4 kbps and of G.723.1 at 5.3 kbps, and it is slightly better than that of G.723.1 at 6.3 kbps.

Greg Pottie - One of the best experts on this subject based on the ideXlab platform.

  • a perceptually based embedded subband Speech Coder
    IEEE Transactions on Speech and Audio Processing, 1997
    Co-Authors: B Tang, Abeer Alwan, A Shen, Greg Pottie
    Abstract:

    A new scheme for robust, high-quality, embedded Speech coding based on subband decomposition and perceptually optimized bit allocation and prioritization is presented. An infinite impulse response (IIR) quadrature mirror filterbank (QMF) performs subband decomposition. A perceptual model, computed using subband spectral analysis, optimizes the Coder's perceptual quality. Dynamic bit allocation and prioritization is combined with embedded quantization resulting in little performance degradation relative to a nonembedded implementation. The Coder output is scalable from high quality at higher bit rates to lower quality at lower bit rates, supporting a wide range of service and resource utilization. The lower bit-rate representation is obtained simply through truncation of the higher bit-rate representation. Since source-rate adaptation is performed through truncation of the encoded stream, interaction with the Coder is not required, making the embedded Coder ideally suited for rate-adaptive communication systems. Performance for both Speech and music was verified through subjective listening tests.

  • a robust variable rate Speech Coder
    International Conference on Acoustics Speech and Signal Processing, 1995
    Co-Authors: I Shen, B Tang, Abeer Alwan, Greg Pottie
    Abstract:

    The goal of this study is to develop a robust and high-quality Speech Coder for wireless communication. The proposed Coder is a perceptually-based variable-rate subband Coder. The perceptual metric ensures that encoding is optimized to the human listener and is based on calculating the signal-to-mask ratio in short-time frames of the input signal. An adaptive bit allocation scheme is employed and the subband energies are then quantized using a Max-Lloyd quantizer. The Coder is fully scalable-increasing the bit rates, improves the quality of encoded Speech. Subjective listening tests, using quiet and noisy input signals, indicate that the proposed Coder produces high-quality Speech when operating at 12 kbps or higher. In error-free conditions, our Coder has comparable performance to that of QCELP or GSM Coders. For Speech in background noise, however, our Coder, at 12 kbps, outperforms QCELP significantly, and for music, it outperforms both QCELP and GSM.

R V Cox - One of the best experts on this subject based on the ideXlab platform.

  • a bitstream based front end for wireless Speech recognition on is 136 communications system
    IEEE Transactions on Speech and Audio Processing, 2001
    Co-Authors: Hong Kook Kim, R V Cox
    Abstract:

    We propose a feature extraction method for a Speech recognizer that operates in digital communication networks. The feature parameters are basically extracted by converting the quantized spectral information of a Speech Coder into a cepstrum. We also include the voiced/unvoiced information obtained from the bitstream of the Speech Coder in the recognition feature set. We performed speaker-independent connected digit HMM recognition experiments under clean, background noise, and channel impairment conditions. From these results, we found that the Speech recognition system employing the proposed bitstream-based front-end gives superior word and string accuracies over a recognizer constructed from decoded Speech signals. Its performance is comparable to that of a wireline recognition system that uses the cepstrum as a feature set. Next, we extended the evaluation of the proposed bitstream-based front-end to large vocabulary Speech recognition with a name database. The recognition results proved that the proposed bitstream-based front-end also gives a comparable performance to the conventional wireline front-end.

  • a very low bit rate Speech Coder based on a recognition synthesis paradigm
    IEEE Transactions on Speech and Audio Processing, 2001
    Co-Authors: Kiseung Lee, R V Cox
    Abstract:

    Previous studies have shown that a concatenative Speech synthesis system with a large database produces more natural sounding Speech. We apply this paradigm to the design of improved very low bit rate Speech Coders (sub 1000 b/s). The proposed Speech Coder consists of unit selection, prosody coding, prosody modification and waveform concatenation. The enCoder selects the best unit sequence from a large database and compresses the prosody information. The transmitted parameters include unit indices and the prosody information. To increase naturalness as well as intelligibility, two costs are considered in the unit selection process: an acoustic target cost and a concatenation cost. A rate-distortion-based piecewise linear approximation is proposed to compress the pitch contour. The deCoder concatenates the set of units, and then synthesizes the resultant sequence of Speech frames using the harmonic+noise model (HNM) scheme. Before concatenating units, prosody modification which includes pitch shifting and gain modification is applied to match those of the input Speech. With single speaker stimuli, a comparison category rating (CCR) test shows that the performance of the proposed Coder is close to that of the 2400-b/s MELP Coder at an average bit rate of about 800-b/s during talk spurts.

  • tts based very low bit rate Speech Coder
    International Conference on Acoustics Speech and Signal Processing, 1999
    Co-Authors: Kiseung Lee, R V Cox
    Abstract:

    This paper addresses a Speech Coder which uses a text-to-Speech (TTS) synthesis system to achieve very low bit rates (sub 1 kbps). The main issue of the work is the accurate coding of the pitch (f/sub 0/) and gain contours which are principle components of prosody. This is of paramount interest since the correct prosody will increase naturalness and an efficient coding scheme will provide high coding gain. Together with the phonetic transcription, the f/sub 0/ and gain contour constitute the parameters that are necessary for the TTS system to synthesize the Speech signal. Piecewise linear approximation is used to code the f/sub 0/ parameter. A technique which minimizes the bit rate while maintaining f/sub 0/ error below a given threshold are described. To obtain both high compression and smoothly changing gain contours, the variance of the signal is averaged over each half phoneme length is transmitted as gain information. With single speaker stimuli, and a priori text transcription information, we obtained natural sounding Speech at an average bit rate of about 300 bps.

Kiseung Lee - One of the best experts on this subject based on the ideXlab platform.

  • a segmental Speech Coder based on a concatenative tts
    Speech Communication, 2002
    Co-Authors: Kiseung Lee, Richard V Cox
    Abstract:

    An extremely low bit rate Speech Coder based on a recognition/synthesis paradigm is proposed. In our Speech Coder, the Speech signal is produced in a way which is similar to concatenative Speech synthesis of text-to-Speech (TTS). Hence, database construction, unit selection and prosody modification, which are the major parts of concatenative TTS, are employed to implement the Speech Coder. The synthesis units are automatically found in a large database using a joint segmentation/classification scheme. Dynamic programming (DP) is applied to unit selection in which two cost functions, an acoustic target cost and a concatenation cost are used to increase naturalness as well as intelligibility. Prosodic differences between the selected unit and the input segment are compensated for by time-scale and pitch modifications which are based on the harmonic plus noise (HNM) model framework. In single speaker tests, the proposed scheme gave intelligible and natural sounding Speech at an average bit rate of about 580 b/s.

  • a very low bit rate Speech Coder based on a recognition synthesis paradigm
    IEEE Transactions on Speech and Audio Processing, 2001
    Co-Authors: Kiseung Lee, R V Cox
    Abstract:

    Previous studies have shown that a concatenative Speech synthesis system with a large database produces more natural sounding Speech. We apply this paradigm to the design of improved very low bit rate Speech Coders (sub 1000 b/s). The proposed Speech Coder consists of unit selection, prosody coding, prosody modification and waveform concatenation. The enCoder selects the best unit sequence from a large database and compresses the prosody information. The transmitted parameters include unit indices and the prosody information. To increase naturalness as well as intelligibility, two costs are considered in the unit selection process: an acoustic target cost and a concatenation cost. A rate-distortion-based piecewise linear approximation is proposed to compress the pitch contour. The deCoder concatenates the set of units, and then synthesizes the resultant sequence of Speech frames using the harmonic+noise model (HNM) scheme. Before concatenating units, prosody modification which includes pitch shifting and gain modification is applied to match those of the input Speech. With single speaker stimuli, a comparison category rating (CCR) test shows that the performance of the proposed Coder is close to that of the 2400-b/s MELP Coder at an average bit rate of about 800-b/s during talk spurts.

  • tts based very low bit rate Speech Coder
    International Conference on Acoustics Speech and Signal Processing, 1999
    Co-Authors: Kiseung Lee, R V Cox
    Abstract:

    This paper addresses a Speech Coder which uses a text-to-Speech (TTS) synthesis system to achieve very low bit rates (sub 1 kbps). The main issue of the work is the accurate coding of the pitch (f/sub 0/) and gain contours which are principle components of prosody. This is of paramount interest since the correct prosody will increase naturalness and an efficient coding scheme will provide high coding gain. Together with the phonetic transcription, the f/sub 0/ and gain contour constitute the parameters that are necessary for the TTS system to synthesize the Speech signal. Piecewise linear approximation is used to code the f/sub 0/ parameter. A technique which minimizes the bit rate while maintaining f/sub 0/ error below a given threshold are described. To obtain both high compression and smoothly changing gain contours, the variance of the signal is averaged over each half phoneme length is transmitted as gain information. With single speaker stimuli, and a priori text transcription information, we obtained natural sounding Speech at an average bit rate of about 300 bps.

Alan V Mccree - One of the best experts on this subject based on the ideXlab platform.

  • an embedded adaptive multi rate wideband Speech Coder
    International Conference on Acoustics Speech and Signal Processing, 2001
    Co-Authors: Alan V Mccree, Takahiro Unno, A K Anandakumar, A Bernard, Erdal Paksoy
    Abstract:

    This paper presents a multi-rate wideband Speech Coder with bit rates from 8 to 32 kb/s. The Coder uses a splitband approach, where the input signal, sampled at 16 kHz, is split into two equal frequency bands from 0-4 kHz and 4-8 kHz, each of which is decimated to an 8 kHz sampling rate. The lower band is coded using the adaptive multi-rate (AMR) family of high-quality narrowband Speech Coders, while the higher band is represented by a simple but effective parametric model. A complete solution including this wideband Speech Coder, channel coding for various GSM channels, and dynamic rate adaptation, easily passed all Selection Rules and ranked second overall in the 3GPP AMR Wideband Selection Testing. Besides the high performance, additional advantages of the embedded split-band approach include ease of implementation, reduced complexity, and simplified interoperation with narrowband Speech Coders.

  • a 14 kb s wideband Speech Coder with a parametric highband model
    International Conference on Acoustics Speech and Signal Processing, 2000
    Co-Authors: Alan V Mccree
    Abstract:

    This paper describes a new 14 kb/s wideband Speech Coder. The Coder uses a split-band approach, where the input signal, sampled at 16 kHz, is split into two equal frequency bands from 0-4 kHz and 4-8 kHz, each of which is decimated to an 8 kHz sampling rate. The lower band is coded with a high-quality narrowband Speech Coder, the 11.8 kb/s G.729 Annex E, while the higher band is represented by a simple but effective parametric model. Two new features facilitate efficient coding of the high-band signal: noise modulation and high-frequency reversal. Since the encoding of the lower band is independent of the high-band signal, the narrowband enCoder output can be embedded in the overall bitstream. Subjective test results show that this wideband Speech Coder is capable of producing high quality output Speech.

  • an adaptive multi rate Speech Coder for digital cellular telephony
    International Conference on Acoustics Speech and Signal Processing, 1999
    Co-Authors: Erdal Paksoy, Alan V Mccree, A K Anandakumar, Carlos J De Martin, C G Gerlach, Waiming Lai, Vishu R Viswanathan
    Abstract:

    We have developed an adaptive multi-rate (AMR) Speech Coder designed to operate under the GSM digital cellular full rate (22.8 kb/s) and half rate (11.4 kb/s) channels and to maintain high quality in the presence of highly varying background noise and channel conditions. Within each total rate, several codec modes with different source/channel bit rate allocations are used. The Speech Coders in each codec mode are based on the CELP algorithm operating at rates ranging from 11.85 kb/s down to 5.15 kb/s, where the lowest rate Coder is a source controlled multi-modal Speech Coder. The deCoders monitor the channel quality at both ends of the wireless link using the soft values for the received bits and assist the base station in selecting the codec mode that is appropriate for a given channel condition. The Coder was submitted to the GSM AMR standardization competition and met the qualification requirements in an independent formal MOS test.

  • a variable rate multimodal Speech Coder with gain matched analysis by synthesis
    International Conference on Acoustics Speech and Signal Processing, 1997
    Co-Authors: Erdal Paksoy, Alan V Mccree, Vishu R Viswanathan
    Abstract:

    In general, a variable rate Coder can obtain the same Speech quality as a fixed rate Coder, while reducing the average bit rate. We have developed a variable-rate multimodal Speech Coder with an average bit rate of 3 kb/s for a Speech activity factor of 80% and quality comparable to the GSM full rate Coder. The Coder has four coding modes and uses a robust classification method involving the pitch gain, zero crossings, and a peakiness measure. Also the Coder employs a novel gain-matched analysis-by-synthesis technique for very low rate coding of unvoiced frames and an improved noise-level-dependent postfilter. This paper describes the details of our algorithm and presents the results from subjective listening tests.