The Experts below are selected from a list of 150879 Experts worldwide ranked by ideXlab platform
Akihiro Yamamoto - One of the best experts on this subject based on the ideXlab platform.
-
two step physical register deallocation for data prefetching and address pre calculation
Ipsj Online Transactions, 2008Co-Authors: Akihiro Yamamoto, Hideki Ando, Yusuke Tanaka, Toshio ShimadaAbstract:This paper proposes an instruction pre-execution scheme for a high performance Processor, that reduces latency and early scheduling of loads. Our scheme exploits the difference between the amount of instruction-level parallelism available with an unlimited number of physical registers and that available with an actual number of physical registers. We introduce the two-step physical register deallocation scheme, which deallocates physical registers at the renaming stage as a first step, and eliminates pipeline stalls caused by a shortage of physical registers. Instructions wait for the final deallocation as a second step in the instruction window. While waiting, the scheme allows pre-execution of instructions, that enables prefetching of load data and early calculation of memory effective addresses. Our evaluation results show that our scheme improves the performance significantly, and achieves a 1.26 times speedup over a Processor without a prefetcher. If combined with a stride prefetcher, it achieves a 1.18 times speedup over a Processor with a stride prefetcher.
-
data prefetching and address pre calculation through instruction pre execution with two step physical register deallocation
Memory Performance: Dealing With Applications Systems And Architecture, 2007Co-Authors: Akihiro Yamamoto, Hideki Ando, Yusuke Tanaka, Toshio ShimadaAbstract:This paper proposes an instruction pre-execution scheme that reduces latency and early scheduling of loads for a high performance Processor. Our scheme exploits the difference between the available amount of instruction-level parallelism with an unlimited number of physical registers and that with an actual number of physical registers. We introduce a scheme called two-step physical register deallocation. Our scheme deallocates physical registers at the renaming stage as a first step, and eliminates pipeline stalls caused by a physical register shortage. Instructions wait for the final deallocation as a second step in the instruction window. While waiting, the scheme allows pre-execution of instructions. This enables prefetching of load data and early calculation of memory effective addresses. In particular, our execution-based scheme has the strength on prefetch of data with an irregular access pattern. Considering the strength of an automatic prefetcher for a regular access pattern, combining it with our scheme offers the best use of our scheme. The evaluation results show that the combined scheme significantly improve performance over a Processor with an automatic prefetcher.
Toshio Shimada - One of the best experts on this subject based on the ideXlab platform.
-
two step physical register deallocation for data prefetching and address pre calculation
Ipsj Online Transactions, 2008Co-Authors: Akihiro Yamamoto, Hideki Ando, Yusuke Tanaka, Toshio ShimadaAbstract:This paper proposes an instruction pre-execution scheme for a high performance Processor, that reduces latency and early scheduling of loads. Our scheme exploits the difference between the amount of instruction-level parallelism available with an unlimited number of physical registers and that available with an actual number of physical registers. We introduce the two-step physical register deallocation scheme, which deallocates physical registers at the renaming stage as a first step, and eliminates pipeline stalls caused by a shortage of physical registers. Instructions wait for the final deallocation as a second step in the instruction window. While waiting, the scheme allows pre-execution of instructions, that enables prefetching of load data and early calculation of memory effective addresses. Our evaluation results show that our scheme improves the performance significantly, and achieves a 1.26 times speedup over a Processor without a prefetcher. If combined with a stride prefetcher, it achieves a 1.18 times speedup over a Processor with a stride prefetcher.
-
data prefetching and address pre calculation through instruction pre execution with two step physical register deallocation
Memory Performance: Dealing With Applications Systems And Architecture, 2007Co-Authors: Akihiro Yamamoto, Hideki Ando, Yusuke Tanaka, Toshio ShimadaAbstract:This paper proposes an instruction pre-execution scheme that reduces latency and early scheduling of loads for a high performance Processor. Our scheme exploits the difference between the available amount of instruction-level parallelism with an unlimited number of physical registers and that with an actual number of physical registers. We introduce a scheme called two-step physical register deallocation. Our scheme deallocates physical registers at the renaming stage as a first step, and eliminates pipeline stalls caused by a physical register shortage. Instructions wait for the final deallocation as a second step in the instruction window. While waiting, the scheme allows pre-execution of instructions. This enables prefetching of load data and early calculation of memory effective addresses. In particular, our execution-based scheme has the strength on prefetch of data with an irregular access pattern. Considering the strength of an automatic prefetcher for a regular access pattern, combining it with our scheme offers the best use of our scheme. The evaluation results show that the combined scheme significantly improve performance over a Processor with an automatic prefetcher.
S. Bateman - One of the best experts on this subject based on the ideXlab platform.
-
A High-Performance Processor for embedded real-time control
IEEE Transactions on Control Systems Technology, 2005Co-Authors: R. Cumplido, Simon W. Jones, Roger M. Goodall, S. BatemanAbstract:This brief reports on an algorithm and corresponding Processor architecture for the construction of High-Performance Processors targeted at linear time invariant (LTI) control. The overall approach involves reformulating the controller into a particular discrete state-space representation, which is optimized for numerical efficiency using the /spl delta/ operator, then programming this into a specially-designed control system Processor (CSP) implemented using a "programmable ASIC" device. This architecture presents large cost and performance benefits for control applications over traditional architectures, particularly for large multiple-input-multiple-output (MIMO) controllers. Results of implementing control of the vertical modes of a Maglev vehicle are presented and compared with implementations using commercial Processors.
David Blaauw - One of the best experts on this subject based on the ideXlab platform.
-
accurate crosstalk noise modeling for early signal integrity analysis
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2003Co-Authors: Li Ding, David Blaauw, Pinaki MazumderAbstract:In this paper, we propose an accurate and fast method to estimate the crosstalk noise in the presence of multiple aggressor nets for use in physical design automation tools. Since noise estimation is often part of the inner loop of optimization algorithms, very efficient closed-form solutions are needed. Previous approaches model aggressor nets one at a time, assuming that the coupling capacitance to all quiet aggressor nets are grounded. They also model the load from interconnect branches as a lumped capacitor, the value of which is the sum of interconnect and load capacitances of the branch. Finally, previous works typically use simple lumped 2-4-node circuit templates and employ a so-called dominant pole approximation to solve the template circuit. While these approximations allow for very fast analysis, they may result in significant underestimation of the noise. In this paper, we propose a new and more comprehensive fast noise estimation method. We propose a novel reduction technique for modeling quiet aggressor nets based on the concept of coupling point admittance. We also propose a reduction method to replace tree branches with effective capacitors which models the effect of resistive shielding. Furthermore, we model the simplified single aggressor net crosstalk noise problem using a 6-node template circuit and propose a new double pole approach to solve the template circuit. We have tested the proposed method on noise-prone interconnects from an industrial High-Performance Processor. Our results show a worst case error of 7.8% and an average error of 2.7%, while allowing for very fast analysis.
-
clarinet a noise analysis tool for deep submicron design
Design Automation Conference, 2000Co-Authors: Rafi Levy, David Blaauw, Gabi Braca, Aurobindo Dasgupta, Amir Grinshpon, Chanhee Oh, Boaz Orshav, Supamas Sirichotiyakul, Vladimir ZolotovAbstract:Coupled noise analysis has become a critical issue for deep-submicron, high performance design. In this paper, we present, ClariNet, an industrial noise analysis tool, which was developed to efficiently analyze large, high performance Processor designs. We present the overall approach and tool flow of ClariNet and discuss three critical large-Processor design issues which have received limited discussion in the past. First, we present how the driver gates of a coupled interconnect network are represented with accurate linear models. Second, we show how to speed the analysis of large designs by using noise filters based on reduced interconnect representations and then pruning the nets coupled to a signal net. Third, we show how to incorporate logic and timing correlations into noise analysis to reduce its pessimism. We present the results from several industrial circuits, including a large high performance microProcessor design and a DSP design.
-
Power management issues in high performance Processor design
Proceedings IEEE Alessandro Volta Memorial Workshop on Low-Power Design, 1999Co-Authors: David BlaauwAbstract:Summary form only given. With the growing demand for portable applications, low power Processor design is increasingly common. In addition to power requirements, Processors also have very stringent performance requirements. These conflicting goals present the designer with a challenging problem. In order to effectively reach an optimal trade-off between performance and power, a number of mature design methods and design tools are needed. The most prominent and mature low power design tool is a power simulator. Power simulation can be performed at the transistor level, gate level, or RTL level. Each additional level of abstraction increases the performance of the tool but reduces the accuracy of the power estimate. The drive for lower power, as well as process shrink, have led to aggressive reductions in the supply voltage. As a result, to maintain performance, the current needed to supply the chip with power is increasing. Due to the resistance of the interconnect, a small voltage drop develops as the power grid supplies current to the circuitry on the chip. Since the current drawn by the devices fluctuates with time, the voltage delivered to the devices fluctuates. The voltage drop and voltage fluctuation results in a number of problems, such as degraded or unreliable performance, noise injection into the signal lines of the circuit, and electro-migration and reliability concerns. In this presentation, we give an overview of traditional power simulation tools and discussed two emerging power management design technologies: power distribution integrity analysis and standby current measurement and optimization. We present methods for accurate peak current simulation, which is needed for power grid integrity analysis, and discuss the generation and compression of the simulation vectors. Standby leakage current is state dependent and we present methods for calculating both the average and maximum leakage current. Finally, optimization methods for minimizing the leakage current are discussed.
Eugene B. John - One of the best experts on this subject based on the ideXlab platform.
-
Wasted dynamic power and correlation to instruction set architecture for CPU throttling
The Journal of Supercomputing, 2019Co-Authors: Abdullah A. Owahid, Eugene B. JohnAbstract:Reducing dynamic power consumption is one of the major design goals in modern High-Performance Processor design. Throttling is a mechanism that reduces dynamic power at the expense of reduced throughput. Instruction profiling can identify a set of instructions suitable for fine-grained throttling without significant performance degradation. In this paper, an Electronic Design Automation (EDA) flow was developed to process pipeline trace at an early stage to identify the bottleneck. Using the developed EDA flow, this work identifies a set of instructions suitable for fine-grained CPU throttling to reduce wasted dynamic power in RISC-V architecture. To rank higher stall causing instructions in the instruction profile, a weight-based system was introduced. It was observed that independent of the workload and type, higher stall causing instructions were repeating across all the benchmark programs. The top 10 instruction profiles for each test suite identify probable throttling clock cycles for each pipeline stage for wasted dynamic power reduction at minimal performance loss. These results are expected to enable researchers to reduce wasted dynamic power by modifying existing architecture and effectively apply throttling mechanism without significant performance degradation.
-
Wasted dynamic power and correlation to instruction set architecture for CPU throttling
The Journal of Supercomputing, 2019Co-Authors: Abdullah A. Owahid, Eugene B. JohnAbstract:Reducing dynamic power consumption is one of the major design goals in modern High-Performance Processor design. Throttling is a mechanism that reduces dynamic power at the expense of reduced throughput. Instruction profiling can identify a set of instructions suitable for fine-grained throttling without significant performance degradation. In this paper, an Electronic Design Automation (EDA) flow was developed to process pipeline trace at an early stage to identify the bottleneck. Using the developed EDA flow, this work identifies a set of instructions suitable for fine-grained CPU throttling to reduce wasted dynamic power in RISC-V architecture. To rank higher stall causing instructions in the instruction profile, a weight-based system was introduced. It was observed that independent of the workload and type, higher stall causing instructions were repeating across all the benchmark programs. The top 10 instruction profiles for each test suite identify probable throttling clock cycles for each pipeline stage for wasted dynamic power reduction at minimal performance loss. These results are expected to enable researchers to reduce wasted dynamic power by modifying existing architecture and effectively apply throttling mechanism without significant performance degradation.