The Experts below are selected from a list of 27588 Experts worldwide ranked by ideXlab platform
Dmitry Ponomarev - One of the best experts on this subject based on the ideXlab platform.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
Oguz Ergin - One of the best experts on this subject based on the ideXlab platform.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
Kanad Ghose - One of the best experts on this subject based on the ideXlab platform.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
Deniz Balkan - One of the best experts on this subject based on the ideXlab platform.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
-
Register packing exploiting narrow width operands for reducing Register file pressure
International Symposium on Microarchitecture, 2004Co-Authors: Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry PonomarevAbstract:A large percentage of computed results have fewer significant bits compared to the full width of a Register. We exploit this fact to pack multiple results into a Single physical Register to reduce the pressure on the Register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a Single Register are evaluated. The first scheme is conservative and allocates a full-width Register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common Register, freeing up the full-width Register. The second scheme allocates Register partitions based on a prediction of the width of the result and reallocates Register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat Register-constrained datapath.
Rajiv Gupta - One of the best experts on this subject based on the ideXlab platform.
-
Bitwidth aware global Register allocation
2003Co-Authors: Sriraman Tallam, Rajiv GuptaAbstract:Multimedia and network processing applications make extensive use of subword data. Since Registers are capable of holding a full data word, when a subword variable is assigned a Register, only part of the Register is used. New embedded processors have started sup-porting instruction sets that allow direct referencing of bit sections within Registers and therefore multiple subword variables can be made to simultaneously reside in the same Register without hinder-ing accesses to these variables. However, a new Register allocation algorithm is needed that is aware of the bitwidths of program vari-ables and is capable of packing multiple subword variables into a Single Register. This paper presents one such algorithm. The algorithm we propose has two key steps. First, a combina-tion of forward and backward data flow analyses are developed to determine the bitwidths of program variables throughout the pro-gram. This analysis is required because the declared bitwidths of variables are often larger than their true bitwidths and moreover the minimal bitwidths of a program variable can vary from one pro-gram point to another. Second, a novel interference graph represen-tation is designed to enable support for a fast and highly accurate algorithm for packing of subword variables into a Single Register. Packing is carried out by a node coalescing phase that precedes the conventional graph coloring phase of Register allocation. In contrast to traditional node coalescing, packing coalesces a set of interfering nodes. Our experiments show that our bitwidth aware Register allo-cation algorithm reduces the Register requirements by 10 % to 50% over a traditional Register allocation algorithm that assigns separate Registers to simultaneously live subword variables
-
Bitwidth Aware Global Register Allocation
ACM, 2002Co-Authors: Sriraman Tallam, Rajiv GuptaAbstract:use of subword data. Since Registers are capable of holding a full data word, when a subword variable is assigned a Register, only part of the Register is used. New embedded processors have started supporting instruction sets that allow direct referencing of bit sections within Registers and therefore multiple subword variables can be made to simultaneously reside in the same Register without hindering accesses to these variables. However, a new Register allocation algorithm is needed that is aware of the bitwidths of program variables and is capable of packing multiple subword variables into a Single Register. This paper presents one such algorithm