The Experts below are selected from a list of 75 Experts worldwide ranked by ideXlab platform
Won Woo Ro - One of the best experts on this subject based on the ideXlab platform.
-
Linebacker: Preserving Victim Cache Lines in Idle Register Files of GPUs
2019 ACM IEEE 46th Annual International Symposium on Computer Architecture (ISCA), 2019Co-Authors: Yunho Oh, Murali Annavaram, Won Woo RoAbstract:Modern GPUs suffer from cache contention due to the limited cache size that is shared across tens of concurrently running warps. To increase the per-warp cache size prior techniques proposed warp throttling which limits the number of active warps. Warp throttling leaves several Registers to be dynamically unused whenever a warp is throttled. Given the stringent cache size limitation in GPUs this work proposes a new cache management technique named Linebacker (LB) that improves GPU performance by utilizing idle Register file space as victim cache space. Whenever a CTA becomes inactive, linebacker backs up the Registers of the throttled CTA to the off-chip memory. Then, linebacker utilizes the corresponding Register file space as victim cache space. If any load instruction finds data in the victim cache line, the data is directly copied to the Destination Register through a simple Register-Register move operation. To further improve the efficiency of victim cache linebacker allocates victim cache space only to a select few load instructions that exhibit high data locality. Through a careful design of victim cache indexing and management scheme linebacker provides 29.0% of speedup compared to the previously proposed warp throttling techniques.
-
ISCA - Linebacker: preserving victim cache lines in idle Register files of GPUs
Proceedings of the 46th International Symposium on Computer Architecture, 2019Co-Authors: Yunho Oh, Murali Annavaram, Won Woo RoAbstract:Modern GPUs suffer from cache contention due to the limited cache size that is shared across tens of concurrently running warps. To increase the per-warp cache size prior techniques proposed warp throttling which limits the number of active warps. Warp throttling leaves several Registers to be dynamically unused whenever a warp is throttled. Given the stringent cache size limitation in GPUs this work proposes a new cache management technique named Linebacker (LB) that improves GPU performance by utilizing idle Register file space as victim cache space. Whenever a CTA becomes inactive, linebacker backs up the Registers of the throttled CTA to the off-chip memory. Then, linebacker utilizes the corresponding Register file space as victim cache space. If any load instruction finds data in the victim cache line, the data is directly copied to the Destination Register through a simple Register-Register move operation. To further improve the efficiency of victim cache linebacker allocates victim cache space only to a select few load instructions that exhibit high data locality. Through a careful design of victim cache indexing and management scheme linebacker provides 29.0% of speedup compared to the previously proposed warp throttling techniques.
Yunho Oh - One of the best experts on this subject based on the ideXlab platform.
-
Linebacker: Preserving Victim Cache Lines in Idle Register Files of GPUs
2019 ACM IEEE 46th Annual International Symposium on Computer Architecture (ISCA), 2019Co-Authors: Yunho Oh, Murali Annavaram, Won Woo RoAbstract:Modern GPUs suffer from cache contention due to the limited cache size that is shared across tens of concurrently running warps. To increase the per-warp cache size prior techniques proposed warp throttling which limits the number of active warps. Warp throttling leaves several Registers to be dynamically unused whenever a warp is throttled. Given the stringent cache size limitation in GPUs this work proposes a new cache management technique named Linebacker (LB) that improves GPU performance by utilizing idle Register file space as victim cache space. Whenever a CTA becomes inactive, linebacker backs up the Registers of the throttled CTA to the off-chip memory. Then, linebacker utilizes the corresponding Register file space as victim cache space. If any load instruction finds data in the victim cache line, the data is directly copied to the Destination Register through a simple Register-Register move operation. To further improve the efficiency of victim cache linebacker allocates victim cache space only to a select few load instructions that exhibit high data locality. Through a careful design of victim cache indexing and management scheme linebacker provides 29.0% of speedup compared to the previously proposed warp throttling techniques.
-
ISCA - Linebacker: preserving victim cache lines in idle Register files of GPUs
Proceedings of the 46th International Symposium on Computer Architecture, 2019Co-Authors: Yunho Oh, Murali Annavaram, Won Woo RoAbstract:Modern GPUs suffer from cache contention due to the limited cache size that is shared across tens of concurrently running warps. To increase the per-warp cache size prior techniques proposed warp throttling which limits the number of active warps. Warp throttling leaves several Registers to be dynamically unused whenever a warp is throttled. Given the stringent cache size limitation in GPUs this work proposes a new cache management technique named Linebacker (LB) that improves GPU performance by utilizing idle Register file space as victim cache space. Whenever a CTA becomes inactive, linebacker backs up the Registers of the throttled CTA to the off-chip memory. Then, linebacker utilizes the corresponding Register file space as victim cache space. If any load instruction finds data in the victim cache line, the data is directly copied to the Destination Register through a simple Register-Register move operation. To further improve the efficiency of victim cache linebacker allocates victim cache space only to a select few load instructions that exhibit high data locality. Through a careful design of victim cache indexing and management scheme linebacker provides 29.0% of speedup compared to the previously proposed warp throttling techniques.
Murali Annavaram - One of the best experts on this subject based on the ideXlab platform.
-
Linebacker: Preserving Victim Cache Lines in Idle Register Files of GPUs
2019 ACM IEEE 46th Annual International Symposium on Computer Architecture (ISCA), 2019Co-Authors: Yunho Oh, Murali Annavaram, Won Woo RoAbstract:Modern GPUs suffer from cache contention due to the limited cache size that is shared across tens of concurrently running warps. To increase the per-warp cache size prior techniques proposed warp throttling which limits the number of active warps. Warp throttling leaves several Registers to be dynamically unused whenever a warp is throttled. Given the stringent cache size limitation in GPUs this work proposes a new cache management technique named Linebacker (LB) that improves GPU performance by utilizing idle Register file space as victim cache space. Whenever a CTA becomes inactive, linebacker backs up the Registers of the throttled CTA to the off-chip memory. Then, linebacker utilizes the corresponding Register file space as victim cache space. If any load instruction finds data in the victim cache line, the data is directly copied to the Destination Register through a simple Register-Register move operation. To further improve the efficiency of victim cache linebacker allocates victim cache space only to a select few load instructions that exhibit high data locality. Through a careful design of victim cache indexing and management scheme linebacker provides 29.0% of speedup compared to the previously proposed warp throttling techniques.
-
ISCA - Linebacker: preserving victim cache lines in idle Register files of GPUs
Proceedings of the 46th International Symposium on Computer Architecture, 2019Co-Authors: Yunho Oh, Murali Annavaram, Won Woo RoAbstract:Modern GPUs suffer from cache contention due to the limited cache size that is shared across tens of concurrently running warps. To increase the per-warp cache size prior techniques proposed warp throttling which limits the number of active warps. Warp throttling leaves several Registers to be dynamically unused whenever a warp is throttled. Given the stringent cache size limitation in GPUs this work proposes a new cache management technique named Linebacker (LB) that improves GPU performance by utilizing idle Register file space as victim cache space. Whenever a CTA becomes inactive, linebacker backs up the Registers of the throttled CTA to the off-chip memory. Then, linebacker utilizes the corresponding Register file space as victim cache space. If any load instruction finds data in the victim cache line, the data is directly copied to the Destination Register through a simple Register-Register move operation. To further improve the efficiency of victim cache linebacker allocates victim cache space only to a select few load instructions that exhibit high data locality. Through a careful design of victim cache indexing and management scheme linebacker provides 29.0% of speedup compared to the previously proposed warp throttling techniques.
Antonio González - One of the best experts on this subject based on the ideXlab platform.
-
HPCA - A Novel Register Renaming Technique for Out-of-Order Processors
2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2018Co-Authors: Hamid Tabani, Jose-maria Arnau, Jordi Tubella, Antonio GonzálezAbstract:Modern superscalar processors support a large number of in-flight instructions, which requires sizeable Register files. Conventional Register renaming techniques allocate a new storage location, i.e. physical Register, for every instruction whose Destination is a logical Register in order to remove false dependences. Physical Registers are released in a conservative manner when the same logical Register is redefined. For this reason, many cycles may happen between the last read and the release of a physical Register, leading to suboptimal utilization of the Register file. We have observed that for more than 50% of the instructions in SPECfp and more than 30% of the instructions in SPECint that have a Destination Register, the produced value has only a single consumer. In this case, the RAW dependence guarantees that the producer-consumer instructions pair will be executed in program order and, hence, the same physical Register can be used to store the value produced by both instructions. In this paper, we propose a renaming technique that exploits this property to reduce the pressure on the Register file. Our technique leverages physical Register sharing by introducing minor changes in the Register map table and the issue queue. We also describe how our renaming scheme supports precise exceptions. We evaluated our renaming technique on top of a modern out-of-order processor. Our experimental results show that it provides 6% speedup on average for the SPEC2006 benchmarks. Alternatively, our renaming scheme achieves the same performance while reducing the number of physical Registers by 10.5%.
-
A Novel Register Renaming Technique for Out-of-Order Processors
2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2018Co-Authors: Hamid Tabani, Jose-maria Arnau, Jordi Tubella, Antonio GonzálezAbstract:Modern superscalar processors support a large number of in-flight instructions, which requires sizeable Register files. Conventional Register renaming techniques allocate a new storage location, i.e. physical Register, for every instruction whose Destination is a logical Register in order to remove false dependences. Physical Registers are released in a conservative manner when the same logical Register is redefined. For this reason, many cycles may happen between the last read and the release of a physical Register, leading to suboptimal utilization of the Register file. We have observed that for more than 50% of the instructions in SPECfp and more than 30% of the instructions in SPECint that have a Destination Register, the produced value has only a single consumer. In this case, the RAW dependence guarantees that the producer-consumer instructions pair will be executed in program order and, hence, the same physical Register can be used to store the value produced by both instructions. In this paper, we propose a renaming technique that exploits this property to reduce the pressure on the Register file. Our technique leverages physical Register sharing by introducing minor changes in the Register map table and the issue queue. We also describe how our renaming scheme supports precise exceptions. We evaluated our renaming technique on top of a modern out-of-order processor. Our experimental results show that it provides 6% speedup on average for the SPEC2006 benchmarks. Alternatively, our renaming scheme achieves the same performance while reducing the number of physical Registers by 10.5%.
-
ICS - Selective predicate prediction for out-of-order processors
Proceedings of the 20th annual international conference on Supercomputing - ICS '06, 2006Co-Authors: Eduardo Quiñones, Joan-manuel Parcerisa, Antonio GonzálezAbstract:If-conversion transforms control dependencies to data dependencies by using a predication mechanism. It is useful to eliminate hard-to-predict branches and to reduce the severe performance impact of branch mispredictions. However, the use of predicated execution in out-of-order processors has to deal with two problems: there can be multiple definitions for a single Destination Register at rename time, and instructions with a false predicated consume unnecessary resources. Predicting predicates is an effective approach to address both problems. However, predicting predicates that come from hard-to-predict branches is not beneficial in general, because this approach reverses the if-conversion transformation, loosing its potential benefits. In this paper we propose a new scheme that dynamically selects which predicates are worthy to be predicted, and which one are more effective in its if-converted form. We show that our approach significantly outperforms previous proposed schemes. Moreover it performs within 5% of an ideal scheme with perfect predicate prediction.
Hamid Tabani - One of the best experts on this subject based on the ideXlab platform.
-
HPCA - A Novel Register Renaming Technique for Out-of-Order Processors
2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2018Co-Authors: Hamid Tabani, Jose-maria Arnau, Jordi Tubella, Antonio GonzálezAbstract:Modern superscalar processors support a large number of in-flight instructions, which requires sizeable Register files. Conventional Register renaming techniques allocate a new storage location, i.e. physical Register, for every instruction whose Destination is a logical Register in order to remove false dependences. Physical Registers are released in a conservative manner when the same logical Register is redefined. For this reason, many cycles may happen between the last read and the release of a physical Register, leading to suboptimal utilization of the Register file. We have observed that for more than 50% of the instructions in SPECfp and more than 30% of the instructions in SPECint that have a Destination Register, the produced value has only a single consumer. In this case, the RAW dependence guarantees that the producer-consumer instructions pair will be executed in program order and, hence, the same physical Register can be used to store the value produced by both instructions. In this paper, we propose a renaming technique that exploits this property to reduce the pressure on the Register file. Our technique leverages physical Register sharing by introducing minor changes in the Register map table and the issue queue. We also describe how our renaming scheme supports precise exceptions. We evaluated our renaming technique on top of a modern out-of-order processor. Our experimental results show that it provides 6% speedup on average for the SPEC2006 benchmarks. Alternatively, our renaming scheme achieves the same performance while reducing the number of physical Registers by 10.5%.
-
A Novel Register Renaming Technique for Out-of-Order Processors
2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2018Co-Authors: Hamid Tabani, Jose-maria Arnau, Jordi Tubella, Antonio GonzálezAbstract:Modern superscalar processors support a large number of in-flight instructions, which requires sizeable Register files. Conventional Register renaming techniques allocate a new storage location, i.e. physical Register, for every instruction whose Destination is a logical Register in order to remove false dependences. Physical Registers are released in a conservative manner when the same logical Register is redefined. For this reason, many cycles may happen between the last read and the release of a physical Register, leading to suboptimal utilization of the Register file. We have observed that for more than 50% of the instructions in SPECfp and more than 30% of the instructions in SPECint that have a Destination Register, the produced value has only a single consumer. In this case, the RAW dependence guarantees that the producer-consumer instructions pair will be executed in program order and, hence, the same physical Register can be used to store the value produced by both instructions. In this paper, we propose a renaming technique that exploits this property to reduce the pressure on the Register file. Our technique leverages physical Register sharing by introducing minor changes in the Register map table and the issue queue. We also describe how our renaming scheme supports precise exceptions. We evaluated our renaming technique on top of a modern out-of-order processor. Our experimental results show that it provides 6% speedup on average for the SPEC2006 benchmarks. Alternatively, our renaming scheme achieves the same performance while reducing the number of physical Registers by 10.5%.