The Experts below are selected from a list of 12720 Experts worldwide ranked by ideXlab platform
Pavel Zemcik - One of the best experts on this subject based on the ideXlab platform.
-
minimum memory vectorisation of wavelet lifting
Advanced Concepts for Intelligent Vision Systems, 2013Co-Authors: David Barina, Pavel ZemcikAbstract:With the start of the widespread use of discrete wavelet transform the need for its effective implementation is becoming increasingly more important. This work presents a novel approach to discrete wavelet transform through a new computational scheme of wavelet lifting. The presented approach is compared with two other. The results are obtained on a general purpose processor with 4-fold SIMD instruction set (such as Intel x86-64 processors). Using the frequently exploited CDF 9/7 wavelet, the achieved speedup is about 3× compared to naive implementation.
-
ACIVS - Minimum Memory Vectorisation of Wavelet Lifting
Advanced Concepts for Intelligent Vision Systems, 2013Co-Authors: David Barina, Pavel ZemcikAbstract:With the start of the widespread use of discrete wavelet transform the need for its effective implementation is becoming increasingly more important. This work presents a novel approach to discrete wavelet transform through a new computational scheme of wavelet lifting. The presented approach is compared with two other. The results are obtained on a general purpose processor with 4-fold SIMD instruction set (such as Intel x86-64 processors). Using the frequently exploited CDF 9/7 wavelet, the achieved speedup is about 3× compared to naive implementation.
He Zhiqiang - One of the best experts on this subject based on the ideXlab platform.
-
analysis for singal processing development with general purpose processor
International Conference on Communications, 2012Co-Authors: He Zhiqiang, Sun Jianxing, Duan Ran, Yue ChenAbstract:This paper presents an analysis for implementation of high-throughput wireless communication signal processing system on general purpose processor (GPP). The large amount of complex computation brings challenge for the conventional platforms. GPP was compared with DSP and FPGA, and regarded as a new direction for signal processing system development. The program optimization on GPP platforms is necessary based on our research and developing experience. Therefore, we come up with the basic principle and optimization methods for GPP realization, especially for the application of multicore technique. Our analysis is proved to be convincible by solid justifications shown in the paper.
-
major optimization methods for td lte signal processing based on general purpose processor
International Conference on Communications, 2012Co-Authors: Jiang Weipeng, He Zhiqiang, Duan Ran, Wang XinglinAbstract:In this paper, an overview of optimization methods for signal processing based on general purpose processor (GPP) is presented. It includes look-up-table (LUT), Single Instruction Multiple Data (SIMD) and Intel Integrated Performance Primitives (Intel IPP) Library. Utilizing these methods, the physical layer process modules in TD-LTE system can achieve huge benefits in real-time performance. The advantages and drawbacks of these optimization methods are analyzed and discussed, which offers a reference for realizing real-time communication systems on GPP.
-
an implementation of mimo detection in td lte based on general purpose processor
International Conference on Communications, 2012Co-Authors: Li Zhou, He Zhiqiang, Duan Ran, He LifengAbstract:This paper presents a method of implementing Multiple Input Multiple Output (MIMO) detection on general purpose processor (GPP) with Single Instruction Multiple Data (SIMD) instructions. In the receiver of Long Term Evolution (LTE) communication system MIMO detection plays a rather important role. How to reduce the processing latency of MIMO detection becomes increasingly significant. The advantages of GPP compared with programmable hardware provide us a considerable solution to cope with this problem. SIMD is another effective manner as it is able to increase the computing parallel degree. For each algorithm of MIMO detection, fixed-point operations which have higher parallel degree than floating-point operations may be a better choice to further reduce the process delay. Nevertheless, the performance attained by fixed-point is still a factor need to be considered while its latency is quite low.
-
CHINACOM - Analysis for singal processing development with general purpose processor
7th International Conference on Communications and Networking in China, 2012Co-Authors: He Zhiqiang, Sun Jianxing, Duan Ran, Yue ChenAbstract:This paper presents an analysis for implementation of high-throughput wireless communication signal processing system on general purpose processor (GPP). The large amount of complex computation brings challenge for the conventional platforms. GPP was compared with DSP and FPGA, and regarded as a new direction for signal processing system development. The program optimization on GPP platforms is necessary based on our research and developing experience. Therefore, we come up with the basic principle and optimization methods for GPP realization, especially for the application of multicore technique. Our analysis is proved to be convincible by solid justifications shown in the paper.
-
CHINACOM - Major optimization methods for TD-LTE signal processing based on general purpose processor
7th International Conference on Communications and Networking in China, 2012Co-Authors: Jiang Weipeng, He Zhiqiang, Duan Ran, Wang XinglinAbstract:In this paper, an overview of optimization methods for signal processing based on general purpose processor (GPP) is presented. It includes look-up-table (LUT), Single Instruction Multiple Data (SIMD) and Intel Integrated Performance Primitives (Intel IPP) Library. Utilizing these methods, the physical layer process modules in TD-LTE system can achieve huge benefits in real-time performance. The advantages and drawbacks of these optimization methods are analyzed and discussed, which offers a reference for realizing real-time communication systems on GPP.
Duan Ran - One of the best experts on this subject based on the ideXlab platform.
-
analysis for singal processing development with general purpose processor
International Conference on Communications, 2012Co-Authors: He Zhiqiang, Sun Jianxing, Duan Ran, Yue ChenAbstract:This paper presents an analysis for implementation of high-throughput wireless communication signal processing system on general purpose processor (GPP). The large amount of complex computation brings challenge for the conventional platforms. GPP was compared with DSP and FPGA, and regarded as a new direction for signal processing system development. The program optimization on GPP platforms is necessary based on our research and developing experience. Therefore, we come up with the basic principle and optimization methods for GPP realization, especially for the application of multicore technique. Our analysis is proved to be convincible by solid justifications shown in the paper.
-
major optimization methods for td lte signal processing based on general purpose processor
International Conference on Communications, 2012Co-Authors: Jiang Weipeng, He Zhiqiang, Duan Ran, Wang XinglinAbstract:In this paper, an overview of optimization methods for signal processing based on general purpose processor (GPP) is presented. It includes look-up-table (LUT), Single Instruction Multiple Data (SIMD) and Intel Integrated Performance Primitives (Intel IPP) Library. Utilizing these methods, the physical layer process modules in TD-LTE system can achieve huge benefits in real-time performance. The advantages and drawbacks of these optimization methods are analyzed and discussed, which offers a reference for realizing real-time communication systems on GPP.
-
an implementation of mimo detection in td lte based on general purpose processor
International Conference on Communications, 2012Co-Authors: Li Zhou, He Zhiqiang, Duan Ran, He LifengAbstract:This paper presents a method of implementing Multiple Input Multiple Output (MIMO) detection on general purpose processor (GPP) with Single Instruction Multiple Data (SIMD) instructions. In the receiver of Long Term Evolution (LTE) communication system MIMO detection plays a rather important role. How to reduce the processing latency of MIMO detection becomes increasingly significant. The advantages of GPP compared with programmable hardware provide us a considerable solution to cope with this problem. SIMD is another effective manner as it is able to increase the computing parallel degree. For each algorithm of MIMO detection, fixed-point operations which have higher parallel degree than floating-point operations may be a better choice to further reduce the process delay. Nevertheless, the performance attained by fixed-point is still a factor need to be considered while its latency is quite low.
-
CHINACOM - Analysis for singal processing development with general purpose processor
7th International Conference on Communications and Networking in China, 2012Co-Authors: He Zhiqiang, Sun Jianxing, Duan Ran, Yue ChenAbstract:This paper presents an analysis for implementation of high-throughput wireless communication signal processing system on general purpose processor (GPP). The large amount of complex computation brings challenge for the conventional platforms. GPP was compared with DSP and FPGA, and regarded as a new direction for signal processing system development. The program optimization on GPP platforms is necessary based on our research and developing experience. Therefore, we come up with the basic principle and optimization methods for GPP realization, especially for the application of multicore technique. Our analysis is proved to be convincible by solid justifications shown in the paper.
-
CHINACOM - Major optimization methods for TD-LTE signal processing based on general purpose processor
7th International Conference on Communications and Networking in China, 2012Co-Authors: Jiang Weipeng, He Zhiqiang, Duan Ran, Wang XinglinAbstract:In this paper, an overview of optimization methods for signal processing based on general purpose processor (GPP) is presented. It includes look-up-table (LUT), Single Instruction Multiple Data (SIMD) and Intel Integrated Performance Primitives (Intel IPP) Library. Utilizing these methods, the physical layer process modules in TD-LTE system can achieve huge benefits in real-time performance. The advantages and drawbacks of these optimization methods are analyzed and discussed, which offers a reference for realizing real-time communication systems on GPP.
David Barina - One of the best experts on this subject based on the ideXlab platform.
-
minimum memory vectorisation of wavelet lifting
Advanced Concepts for Intelligent Vision Systems, 2013Co-Authors: David Barina, Pavel ZemcikAbstract:With the start of the widespread use of discrete wavelet transform the need for its effective implementation is becoming increasingly more important. This work presents a novel approach to discrete wavelet transform through a new computational scheme of wavelet lifting. The presented approach is compared with two other. The results are obtained on a general purpose processor with 4-fold SIMD instruction set (such as Intel x86-64 processors). Using the frequently exploited CDF 9/7 wavelet, the achieved speedup is about 3× compared to naive implementation.
-
ACIVS - Minimum Memory Vectorisation of Wavelet Lifting
Advanced Concepts for Intelligent Vision Systems, 2013Co-Authors: David Barina, Pavel ZemcikAbstract:With the start of the widespread use of discrete wavelet transform the need for its effective implementation is becoming increasingly more important. This work presents a novel approach to discrete wavelet transform through a new computational scheme of wavelet lifting. The presented approach is compared with two other. The results are obtained on a general purpose processor with 4-fold SIMD instruction set (such as Intel x86-64 processors). Using the frequently exploited CDF 9/7 wavelet, the achieved speedup is about 3× compared to naive implementation.
Ville Lappalainen - One of the best experts on this subject based on the ideXlab platform.
-
performance of h 26l video encoder on general purpose processor
Signal Processing Systems, 2003Co-Authors: Ville Lappalainen, Antti Hallapuro, Timo HamalainenAbstract:Two optimized implementations of the emerging ITU-T H.26L video encoder are described. The first, medium-optimized version, is implemented in C and the latter, highly optimized version, utilizes both algorithmic and platform-specific optimizations. Comparisons to a correspondingly optimized H.263/H.263+ implementation are given with the spatial and temporal video quality fixed and the bit rate and complexity varied. On a 733 MHz general-purpose processor, an average encoding speed of 17 frames per second for QCIF sequences is achieved with a 29% reduction in bit rate compared to H.263+. The complexity of H.26L is about 3.4 times more than that of H.263+.
-
Complexity of optimized H.26L video decoder implementation
IEEE Transactions on Circuits and Systems for Video Technology, 2003Co-Authors: Ville Lappalainen, Antti Hallapuro, Timo D. HämäläinenAbstract:An analysis of computational complexity is presented for an H.26L video decoder, based on extensive experiments on a general-purpose processor. In addition, platform-independent techniques to optimize an H.26L decoder implementation are given. Comparisons are carried out between our highly optimized version of H.26L, the public reference implementation of H.26L, and a highly optimized H.263+ implementation. Both QCIF and CIF-sized image sequences are used. The results show that with equal visual quality, the bit-rate savings range from 28% to 58%, while the frame decoding speed of H.26L is about 11% better than that of a highly optimized H.263+.
-
performance of an advanced video codec on a general purpose processor with media isa extensions
International Conference on Consumer Electronics, 2000Co-Authors: Ville Lappalainen, P Defee, A HallapuroAbstract:This paper analyses the performance of the state-of-the-art media ISA (instruction set architecture) extensions in a general-purpose processor, when executing a video encoder based on an affine motion model. In addition to SIMD (single instruction multiple data) fixed-point instructions, these ISA extensions include SIMD floating-point instructions, special-purpose SIMD fixed-point instructions, and cacheability control instructions. In this study, eight time-consuming kernels of the video encoder were hand-optimized, using instructions in all four instruction categories of these media ISA extensions (the FLP version). These kernels were also hand-optimized using only SIMD fixed-point ISA extensions, without special-purpose instructions (the FXP version). For the FLP version, this study resulted in an average kernel-level speedup of 1.37X and an application-level speedup of 1.11X, compared to the FXP version, and an application-level speedup of 3.41X, compared to the C version.
-
Performance of H.26L video encoder on general-purpose processor
ICCE. International Conference on Consumer Electronics (IEEE Cat. No.01CH37182), 1Co-Authors: Ville Lappalainen, A. Hailapuro, Timo D. HämäläinenAbstract:An optimized implementation of an H.26L video encoder is presented. Compared to H263, H.26L reduces the output bit rate about 25% at the expense of increased (3.8X) complexity. However, optimizations enable real-time operation on a 733 MHz general-purpose processor.