The Experts below are selected from a list of 969 Experts worldwide ranked by ideXlab platform

Sungkyo Jung - One of the best experts on this subject based on the ideXlab platform.

  • Multichannel Voice Transmission and Storage
    2014
    Co-Authors: Sungkyo Jung, Youngcheol Park, Sung-wan Youn, Kyoung-tae Kim, Dae Hee Youn
    Abstract:

    Dual-rate G.723.1 speech coder has been widely applied to real-time video and teleconferencing applications where re-duced bandwidth and good voice quality is required. This paper presents an efficient implementation of G.723.1 speech coder. To simplify the excitation quantization procedure which is the most computationally demanding, we propose fast algorithms for Adaptive Codebook and fixed Codebook search. In the fast Adaptive Codebook search, pitch delay and pitch gains are computed sequentially. In the fast fixed code-book search, the Codebook structure is redesigned based on the interleaved single-pulse permutation (ISPP) design at high rate mode and the depth-first tree search is applied instead of nested-loop search at low rate mode. A real-time implementa-tion is achieved using a 16-bit fixed-point TMS320C62x DSP. The implemented G.723.1 speech coder requires 8.70 and 10.29 MHz clock cycles at low and high rate, respectively, 57.8 kByte of program memory and 55 kByte of data memory. Thus, more than 16 channels of G.723.1 coder can be oper-ated in real-time using a single TMS320C62x DSP. 1

  • Applying a Speaker-Dependent Speech Compression Technique to Concatenative TTS Synthesizers
    IEEE Transactions on Audio Speech and Language Processing, 2007
    Co-Authors: Sungkyo Jung, Honggoo Kang
    Abstract:

    This paper proposes a new speaker-dependent coding algorithm to efficiently compress a large speech database for corpus-based concatenative text-to-speech (TTS) engines while maintaining high fidelity. To achieve a high compression ratio and meet the fundamental requirements of concatenative TTS synthesizers, such as partial segment decoding and random access capability, we adopt a nonpredictive analysis-by-synthesis scheme for speaker-dependent parameter estimation and quantization. The spectral coefficients are quantized by using a memoryless split vector quantization (VQ) approach that does not use frame correlation. Considering that excitation signals of a specific speaker show low intra-variation especially in the voiced regions, the conventional Adaptive Codebook for pitch prediction is replaced by a speaker-dependent pitch-pulse Codebook trained by a corpus of single-speaker speech signals. To further improve the coding efficiency, the proposed coder flexibly combines nonpredictive and predictive type method considering the structure of the TTS system. By applying the proposed algorithm to a Korean TTS system, we could obtain comparable quality to the G.729 speech coder and satisfy all the requirements that TTS system needs. The results are verified by both objective and subjective quality measurements. In addition, the decoding complexity of the proposed coder is around 55% lower than that of G.729 annex A

  • a fast Adaptive Codebook search algorithm for g 723 1 speech coder
    IEEE Signal Processing Letters, 2005
    Co-Authors: Sungkyo Jung, Youngcheol Park, Kyungtae Kim, Honggoo Kang
    Abstract:

    This letter presents a new fast search algorithm for the multitap Adaptive Codebook used in the G.723.1 standard speech coder. In contrast with the standard method that a closed-loop pitch lag and gains for a fifth-order pitch predictor are searched simultaneously, the proposed algorithm adopts a sequential and restricted approach to determine the parameters. In other words, the proposed scheme first determines a couple of pitch lag candidates using a first-order pitch predictor and then computes the pitch gains of the fifth-order predictor within a restricted search area. Experimental results confirm that the proposed algorithm reduces the total complexity by 30.69% in the encoding process and provides speech quality equivalent to the standard method.

  • efficient implementation of itu t g 723 1 speech coder for multichannel voice transmission and storage
    Conference of the International Speech Communication Association, 2001
    Co-Authors: Sungkyo Jung, Youngcheol Park, Sungwan Yoon, Kyungtae Kim, Dae Hee Youn
    Abstract:

    Dual-rate G.723.1 speech coder has been widely applied to real-time video and teleconferencing applications where reduced bandwidth and good voice quality is required. This paper presents an efficient implementation of G.723.1 speech coder. To simplify the excitation quantization procedure which is the most computationally demanding, we propose fast algorithms for Adaptive Codebook and fixed Codebook search. In the fast Adaptive Codebook search, pitch delay and pitch gains are computed sequentially. In the fast fixed Codebook search, the Codebook structure is redesigned based on the interleaved single-pulse permutation (ISPP) design at high rate mode and the depth-first tree search is applied instead of nested-loop search at low rate mode. A real-time implementation is achieved using a 16-bit fixed-point TMS320C62x DSP. The implemented G.723.1 speech coder requires 8.70 and 10.29 MHz clock cycles at low and high rate, respectively, 57.8 kByte of program memory and 55 kByte of data memory. Thus, more than 16 channels of G.723.1 coder can be operated in real-time using a single TMS320C62x DSP.

Suphattharachai Chomphan - One of the best experts on this subject based on the ideXlab platform.

  • Multi-pulse based code excited linear predictive speech coder with fine granularity scalability for tonal language
    2015
    Co-Authors: Suphattharachai Chomphan
    Abstract:

    Abstract: Problem statement: The flexible bit-rate speech coder plays an important role in the modern speech communication. The MP-CELP speech coder which is a candidate of the MPEG4 natural speech coder supports a flexible and wide bit-rate range. However, a fine scalability had not been included. To support finer scalability of the coding rate, it had been studied in this study. Approach: In this study, based on the MP-CELP speech coding with HPDR technique, Fine Granularity Scalability was introduced by adjusting the amount of transmitted fixed excitation information. The FGS feature aim at changing the bit rate of the conventional coding more finely and more smoothly. Results: Through performance analysis and computer simulation, the quality of scalability of the MP-CELP coding was presented with an improvement from conventional scalable MP-CELP. The HPDR technique is also applied to the MP-CELP to use for tonal language, meanwhile it can support the core coding rate of 4.2, 5.5, 7.5 kbps and additional scaled bit rates. Conclusion: The core coder with high pitch delay resolution technique and Adaptive Codebook for tonal speech quality improvement has been conducted and the FGS brings about further efficient scalability. Key words: Flexible bit-rate, speech coder, MP-CELP, fine granularity scalability, bit rate scalability

  • High Pitch Delay Resolution Technique for Tonal Language Speech Coding Based on Multi-Pulse Based Code Excited Linear Prediction Algorithm
    2015
    Co-Authors: Suphattharachai Chomphan
    Abstract:

    Abstract: Problem statement: In spontaneous speech communication, speech coding is an important process that should be taken into account, since the quality of coded speech depends on the efficiency of the speech coding algorithm. As for tonal language which tone plays important role not only on the naturalness and also the intelligibility of the speech, tone must be treated appropriately. Approach: This study proposes a modification of flexible Multi-Pulse based Code Excited Linear Predictive (MP-CELP) coder with multiple bitrates and bitrate scalabilities for tonal language speech in the multimedia applications. The coder consists of a core coder and bitrate scalable tools. The High Pitch Delay Resolutions (HPDR) are applied to the Adaptive Codebook of core coder for tonal language speech quality improvement. The bitrate scalable tool employs multi-stage excitation coding based on an embedded-coding approach. The multi-pulse excitation Codebook at each stage is Adaptively produced depending on the selected excitation signal at the previous stage. Results: The experimental results show that the speech quality of the proposed coder is improved above the speech quality of the conventional coder without pitch-resolution adaptation. Conclusion: From the study, it is a strong evidence to further apply the proposed technique in the speech coding systems or other speech processing technologies

  • Tonal Language Speech Compression Based on a Bitrate Scalable Multi-Pulse Based Code Excited Linear Prediction Coder
    2015
    Co-Authors: Suphattharachai Chomphan
    Abstract:

    Abstract: Problem statement: Speech compression is an important issue in the modern digital speech communication. The functionality of bitrates scalability also plays significant role, since the capacity of communication system varies all the time. When considering tonal speech, such as Thai, tone plays important role on the naturalness and the intelligibility of the speech, it must be treated appropriately. Therefore these issues are taken into account in this study. Approach: This study proposes a modification of flexible Multi-Pulse based Code Excited Linear Predictive (MP-CELP) coder with bitrates scalabilities for tonal language speech in the multimedia applications. The coder consists of a core coder and bitrates scalable tools. The high pitch delay resolutions are applied to the Adaptive Codebook of core coder for tonal language speech quality improvement. The bitrates scalable tool employs multi-stage excitation coding based on an embedded-coding approach. The multi-pulse excitation Codebook at each stage is Adaptively produced depending on the selected excitation signal at the previous stage. Results: The experimental results show that the speech quality of the proposed coder is improved above the speech quality of the conventional coder without pitch-resolution adaptation. Conclusion: From the study, the proposed approach is able to improve the speech compression quality for tonal language and the functionality of bitrates scalability is also developed

  • Multi-Pulse Based Code Excited Linear Predictive Speech Coder with Fine Granularity Scalability for Tonal Language
    2011
    Co-Authors: Suphattharachai Chomphan
    Abstract:

    Abstract: Problem statement: The flexible bit-rate speech coder plays an important role in the modern speech communication. The MP-CELP speech coder which is a candidate of the MPEG4 natural speech coder supports a flexible and wide bit-rate range. However, a fine scalability had not been included. To support finer scalability of the coding rate, it had been studied in this study. Approach: In this study, based on the MP-CELP speech coding with HPDR technique, Fine Granularity Scalability was introduced by adjusting the amount of transmitted fixed excitation information. The FGS feature aim at changing the bit rate of the conventional coding more finely and more smoothly. Results: Through performance analysis and computer simulation, the quality of scalability of the MP-CELP coding was presented with an improvement from conventional scalable MP-CELP. The HPDR technique is also applied to the MP-CELP to use for tonal language, meanwhile it can support the core coding rate of 4.2, 5.5, 7.5 kbps and additional scaled bit rates. Conclusion: The core coder with high pitch delay resolution technique and Adaptive Codebook for tonal speech quality improvement has been conducted and the FGS brings about further efficient scalability. Key words: Flexible bit-rate, speech coder, MP-CELP, fine granularity scalability, bit rate scalability, HPDR techniqu

  • multi pulse based code excited linear predictive speech coder with fine granularity scalability for tonal language
    Journal of Computer Science, 2010
    Co-Authors: Suphattharachai Chomphan
    Abstract:

    Problem statement: The flexible bit-rate speech coder plays an important role in the modern speech communication. The MP-CELP speech coder which is a candidate of the MPEG4 natural speech coder supports a flexible and wide bit-rate range. However, a fine scalability had not been included. To support finer scalability of the coding rate, it had been studied in this study. Approach: In this study, based on the MP-CELP speech coding with HPDR technique, Fine Granularity Scalability was introduced by adjusting the amount of transmitted fixed excitation information. The FGS feature aim at changing the bit rate of the conventional coding more finely and more smoothly. Results: Through performance analysis and computer simulation, the quality of scalability of the MP-CELP coding was presented with an improvement from conventional scalable MP-CELP. The HPDR technique is also applied to the MP-CELP to use for tonal language, meanwhile it can support the core coding rate of 4.2, 5.5, 7.5 kbps and additional scaled bit rates. Conclusion: The core coder with high pitch delay resolution technique and Adaptive Codebook for tonal speech quality improvement has been conducted and the FGS brings about further efficient scalability.

Honggoo Kang - One of the best experts on this subject based on the ideXlab platform.

  • Applying a Speaker-Dependent Speech Compression Technique to Concatenative TTS Synthesizers
    IEEE Transactions on Audio Speech and Language Processing, 2007
    Co-Authors: Sungkyo Jung, Honggoo Kang
    Abstract:

    This paper proposes a new speaker-dependent coding algorithm to efficiently compress a large speech database for corpus-based concatenative text-to-speech (TTS) engines while maintaining high fidelity. To achieve a high compression ratio and meet the fundamental requirements of concatenative TTS synthesizers, such as partial segment decoding and random access capability, we adopt a nonpredictive analysis-by-synthesis scheme for speaker-dependent parameter estimation and quantization. The spectral coefficients are quantized by using a memoryless split vector quantization (VQ) approach that does not use frame correlation. Considering that excitation signals of a specific speaker show low intra-variation especially in the voiced regions, the conventional Adaptive Codebook for pitch prediction is replaced by a speaker-dependent pitch-pulse Codebook trained by a corpus of single-speaker speech signals. To further improve the coding efficiency, the proposed coder flexibly combines nonpredictive and predictive type method considering the structure of the TTS system. By applying the proposed algorithm to a Korean TTS system, we could obtain comparable quality to the G.729 speech coder and satisfy all the requirements that TTS system needs. The results are verified by both objective and subjective quality measurements. In addition, the decoding complexity of the proposed coder is around 55% lower than that of G.729 annex A

  • a fast Adaptive Codebook search algorithm for g 723 1 speech coder
    IEEE Signal Processing Letters, 2005
    Co-Authors: Sungkyo Jung, Youngcheol Park, Kyungtae Kim, Honggoo Kang
    Abstract:

    This letter presents a new fast search algorithm for the multitap Adaptive Codebook used in the G.723.1 standard speech coder. In contrast with the standard method that a closed-loop pitch lag and gains for a fifth-order pitch predictor are searched simultaneously, the proposed algorithm adopts a sequential and restricted approach to determine the parameters. In other words, the proposed scheme first determines a couple of pitch lag candidates using a first-order pitch predictor and then computes the pitch gains of the fifth-order predictor within a restricted search area. Experimental results confirm that the proposed algorithm reduces the total complexity by 30.69% in the encoding process and provides speech quality equivalent to the standard method.

Hsiaochuan Wang - One of the best experts on this subject based on the ideXlab platform.

  • speech classification embedded in Adaptive Codebook search for low bit rate celp coding
    IEEE Transactions on Speech and Audio Processing, 1995
    Co-Authors: Furong Jean, Hsiaochuan Wang
    Abstract:

    This correspondence proposes a new CELP coding method which embeds speech classification in Adaptive Codebook search. This approach can retain the synthesized speech quality at bit-rates below 4 kb/s. A pitch analyzer is designed to classify each frame by its periodicity, and with a finite-state machine, one of four states is determined. Then the Adaptive Codebook search scheme is switched according to the state. Simulation results show that higher SEGSNR and lower computation complexity can be achieved, and the pitch contour of the synthesized speech is smoother than that produced by conventional CELP coders. >

Furong Jean - One of the best experts on this subject based on the ideXlab platform.

  • speech classification embedded in Adaptive Codebook search for low bit rate celp coding
    IEEE Transactions on Speech and Audio Processing, 1995
    Co-Authors: Furong Jean, Hsiaochuan Wang
    Abstract:

    This correspondence proposes a new CELP coding method which embeds speech classification in Adaptive Codebook search. This approach can retain the synthesized speech quality at bit-rates below 4 kb/s. A pitch analyzer is designed to classify each frame by its periodicity, and with a finite-state machine, one of four states is determined. Then the Adaptive Codebook search scheme is switched according to the state. Simulation results show that higher SEGSNR and lower computation complexity can be achieved, and the pitch contour of the synthesized speech is smoother than that produced by conventional CELP coders. >