The Experts below are selected from a list of 10776 Experts worldwide ranked by ideXlab platform
A. M. Fiskiran - One of the best experts on this subject based on the ideXlab platform.
-
plx a fully subword parallel instruction set architecture for fast scalable multimedia processing
International Conference on Multimedia and Expo, 2002Co-Authors: R. B. Lee, A. M. FiskiranAbstract:PLX is a small, fully subword-parallel instruction set architecture (ISA) designed for very fast multimedia processing, especially in constrained environments requiring low cost and power, such as handheld multimedia information appliances. In PLX, we select the most useful multimedia instructions added previously to microprocessors. We also introduce a few novel features: a new definition of predication requiring very few bits in each predicated instruction, and datapath scalability from 32-bit to 128-bit words, which allows different degrees of subword parallelism without any changes to the ISA. Performance results from basic multimedia kernels testify to PLX's superiority for multimedia processing.
-
Refining instruction set architecture for high-performance multimedia processing in constrained environments
Proceedings of the International Conference on Application-Specific Systems Architectures and Processors, 2002Co-Authors: R. B. Lee, A. M. Fiskiran, Zhijie Shi, Xiao YangAbstract:Multimedia processing in software has been significantly accelerated by the addition of subword-parallel instructions to the instruction set architectures (ISAs) of modem microprocessors. While some of these multimedia instructions are simple and effective, others are very complex, requiring large, special-purpose functional units that are not practical for constrained environments such as handheld multimedia information appliances. For such environments, low-power and low-cost are as important as the high performance required for real-time multimedia processing and the general-purpose programmability required to support an ever growing range of applications. In this paper, we introduce PLX, a concise ISA that selects the most useful features from the first two generations of multimedia instructions added to microprocessors, and explores new ISA features for high-performance yet low-cost multimedia processing with small footprint processors. PLX is unique in that it is designed from scratch as a fully subword-parallel architecture with novel features like datapath scalability from 32-bit to 128-bit words, and a new definition of predication for reducing conditional branches. We illustrate the use of PLX's architectural features with four frequently used multimedia kernels: discrete cosine transform, pixel padding, clip test and median filter. Our performance results show that a 64-bit PLX implementation achieves significant speedups compared to a basic 64-bit RISC processor and to IA-32 processors with MMX and SSE multimedia extensions. PLX's datapath scalability feature often provides an additional 2x speedup in a cost-effective way.
R. B. Lee - One of the best experts on this subject based on the ideXlab platform.
-
plx an instruction set architecture and testbed for multimedia information processing
Signal Processing Systems, 2005Co-Authors: R. B. Lee, Murat A FiskiranAbstract:PLX is a concise instruction set architecture (ISA) that combines the most useful features from previous generations of multimedia instruction sets with newer ISA features for high-performance, low-cost multimedia information processing. Unlike previous multimedia instruction sets, PLX is not added onto a base processor ISA, but designed from the beginning as a standalone processor architecture optimized for media processing. Its design goals are high performance multimedia processing, general-purpose programmability to support an ever-growing range of applications, simplicity for constrained environments where low power and low cost are paramount, and scalability for higher performance in less constrained multimedia systems. Another design goal of PLX is to facilitate exploration and evaluation of novel techniques in instruction set architecture, microarchitecture, arithmetic, VLSI implementations, compiler optimizations, and parallel algorithm design for new computing paradigms. Key characteristics of PLX are a fully subword-parallel architecture with novel features like wordsize scalability from 32-bit to 128-bit words, a new definition of predication, and an innovative set of subword permutation instructions. We demonstrate the use and high performance of PLX on some frequently-used code kernels selected from image, video, and graphics processing applications: discrete cosine transform, pixel padding, clip test, and median filter. Our results show that a 64-bit PLX processor achieves significant speedups over a basic 64-bit RISC processor and over IA-32 processors with MMX and SSE multimedia extensions. Using PLX's wordsize scalability feature, PLX-128 often provides an additional 2× speedup over PLX-64 in a cost-effective way. Superscalar or VLIW (Very Long instruction Word) PLX implementations can also add additional performance through inter-instruction, rather than intra-instruction parallelism. We also describe the PLX testbed and its software tools for architecture and related research.
-
plx a fully subword parallel instruction set architecture for fast scalable multimedia processing
International Conference on Multimedia and Expo, 2002Co-Authors: R. B. Lee, A. M. FiskiranAbstract:PLX is a small, fully subword-parallel instruction set architecture (ISA) designed for very fast multimedia processing, especially in constrained environments requiring low cost and power, such as handheld multimedia information appliances. In PLX, we select the most useful multimedia instructions added previously to microprocessors. We also introduce a few novel features: a new definition of predication requiring very few bits in each predicated instruction, and datapath scalability from 32-bit to 128-bit words, which allows different degrees of subword parallelism without any changes to the ISA. Performance results from basic multimedia kernels testify to PLX's superiority for multimedia processing.
-
Refining instruction set architecture for high-performance multimedia processing in constrained environments
Proceedings of the International Conference on Application-Specific Systems Architectures and Processors, 2002Co-Authors: R. B. Lee, A. M. Fiskiran, Zhijie Shi, Xiao YangAbstract:Multimedia processing in software has been significantly accelerated by the addition of subword-parallel instructions to the instruction set architectures (ISAs) of modem microprocessors. While some of these multimedia instructions are simple and effective, others are very complex, requiring large, special-purpose functional units that are not practical for constrained environments such as handheld multimedia information appliances. For such environments, low-power and low-cost are as important as the high performance required for real-time multimedia processing and the general-purpose programmability required to support an ever growing range of applications. In this paper, we introduce PLX, a concise ISA that selects the most useful features from the first two generations of multimedia instructions added to microprocessors, and explores new ISA features for high-performance yet low-cost multimedia processing with small footprint processors. PLX is unique in that it is designed from scratch as a fully subword-parallel architecture with novel features like datapath scalability from 32-bit to 128-bit words, and a new definition of predication for reducing conditional branches. We illustrate the use of PLX's architectural features with four frequently used multimedia kernels: discrete cosine transform, pixel padding, clip test and median filter. Our performance results show that a 64-bit PLX implementation achieves significant speedups compared to a basic 64-bit RISC processor and to IA-32 processors with MMX and SSE multimedia extensions. PLX's datapath scalability feature often provides an additional 2x speedup in a cost-effective way.
Moinul H Khan - One of the best experts on this subject based on the ideXlab platform.
-
intel spl reg wireless mmxtm technology a 64 bit simd architecture for mobile multimedia
International Conference on Acoustics Speech and Signal Processing, 2003Co-Authors: Nigel C Paver, Bradley C Aldrich, Moinul H KhanAbstract:The growing demand for multimedia rich applications in the wireless mobile domain challenges the capabilities of current wireless handheld devices. Optimizing instruction set architecture is a logical approach towards attaining higher performance in multimedia applications. Intel/spl reg/ wireless MMXTM technology is a 64-bit single instruction multiple data, (SIMD), coprocessor for the Intel/spl reg/ XScale/spl trade/ microarchitecture. It accelerates multimedia applications in handheld and wireless devices by taking advantage of the inherent parallelism and data types of targeted applications. This paper provides an overview of the wireless MMX architecture, its instruction set, pipeline organization, and functional units. Initial benchmark results measured on silicon are also presented.
Anhvu Dinhduc - One of the best experts on this subject based on the ideXlab platform.
-
a proposed risc instruction set architecture for the mac unit of 32 bit vliw dsp processor core
International Conference on Computing Management and Telecommunications, 2014Co-Authors: Khoinguyen Lehuu, Anhvu Dinhduc, Quocminh Dangdo, Vy Luu, Trongtu BuiAbstract:Multiplier-accumulator is a specific hardware unit that performs a common operation - computing the product of two numbers and adding that product to an accumulator. Especially, in digital signal processing applications which consist of a large number of convolution operations, the emergence of MAC unit contributes greatly to the high performance of the systems. This work is about an implementation for a specific MAC unit based on the proposed RISC instruction set architecture (ISA) of 32-bit VLIW Fixed-point DSP processor core presented in our previous work. The computational unit is designed to be flexible for 32-bit/16-bit/8-bit data computations. The implementation is verified to function correctly not only in Modelsim software but also on Altera Cyclone II (2C35) FPGA board.
-
towards a risc instruction set architecture for the 32 bit vliw dsp processor core
2014 IEEE REGION 10 SYMPOSIUM, 2014Co-Authors: Khoinguyen Lehuu, Anhvu DinhducAbstract:Digital Signal Processors (DSPs), compared to general-purpose processors, have shown their great contribution to the implementation of digital signal processing algorithms such as digital filtering and Fourier analysis. This work deals with the RISC instruction set architecture (ISA) for the 32-bit VLIW Fixed-point DSP processor core proposed in our previous work. The designed DSP has been described in terms of groups of instructions, the opcode maps, and suggested design of the data path based on the proposed ISA. Moreover, advanced and enhanced instructions aimed at audio and image applications will also be presented in this work.
Andreas Hoffmann - One of the best experts on this subject based on the ideXlab platform.
-
a universal technique for fast and flexible instruction set architecture simulation
Design Automation Conference, 2002Co-Authors: Achim Nohl, Gunnar Braun, Oliver Schliebusch, Rainer Leupers, H Meyr, Andreas HoffmannAbstract:In the last decade, instruction-set simulators have become an essential development tool for the design of new programmable architectures. Consequently, the simulator performance is a key factor for the overall design efficiency. Based on the extremely poor performance of commonly used interpretive simulators, research work on fast compiled instruction-set simulation was started ten years ago. However, due to the restrictiveness of the compiled technique, it has not been able to push through in commercial products. This paper presents a new retargetable simulation technique which combines the performance of traditional compiled simulators with the flexibility of interpretive simulation. This technique is not limited to any class of architectures or applications and can be utilized from architecture exploration up to end-user software development. The work-flow and the applicability of the so-called just-in-time cache compiled simulation (JIT-CCS) technique will be demonstrated by means of state of the art real world architectures.