The Experts below are selected from a list of 33 Experts worldwide ranked by ideXlab platform
William J. Dally - One of the best experts on this subject based on the ideXlab platform.
-
unifying primary cache scratch and register file memories in a Throughput Processor
International Symposium on Microarchitecture, 2012Co-Authors: Mark Gebhart, Stephen W. Keckler, Brucek Khailany, Ronny Krashinsky, William J. DallyAbstract:Modern Throughput Processors such as GPUs employ thousands of threads to drive high-bandwidth, long-latency memory systems. These threads require substantial on-chip storage for registers, cache, and scratchpad memory. Existing designs hard-partition this local storage, fixing the capacities of these structures at design time. We evaluate modern GPU workloads and find that they have widely varying capacity needs across these different functions. Therefore, we propose a unified local memory which can dynamically change the partitioning among registers, cache, and scratchpad on a per-application basis. The tuning that this flexibility enables improves both performance and energy consumption, and broadens the scope of applications that can be efficiently executed on GPUs. Compared to a hard-partitioned design, we show that unified local memory provides a performance benefit as high as 71% along with an energy reduction up to 33%.
-
A Compile-Time Managed Multi-Level Register File Hierarchy
2012Co-Authors: Mark Gebhart, Stephen W. Keckler, William J. DallyAbstract:As Processors increasingly become power limited, performance improvements will be achieved by rearchitecting systems with energy efficiency as the primary design constraint. While some of these optimizations will be hardware based, combined hardware and software techniques likely will be the most productive. This work redesigns the register file system of a modern Throughput Processor with a combined hardware and software solution that reduces register file energy without harming system performance. Throughput Processors utilize a large number of threads to tolerate latency, requiring a large, energy-intensive register file to store thread context. Our results show that a compiler controlled register file hierarchy can reduce register file energy by up to 54%, compared to a hardware only caching approach that reduces register file energy by 34%. We explore register allocation algorithms that are specifically targeted to improve energy efficiency by sharing temporary register file resources across concurrently running threads and conduct a detailed limit study on the further potential to optimize operand delivery for Throughput Processors. Our efficiency gains represent a direct performance gain for power limited systems, such as GPUs
Zibin Dai - One of the best experts on this subject based on the ideXlab platform.
-
a high Throughput Processor for dual field elliptic curve cryptography with power analysis resistance
Ubiquitous Intelligence and Computing, 2015Co-Authors: Xiaoyang Zeng, Xiao Feng, Zibin DaiAbstract:In this paper, a mixture of VLIW and vector architecture of ECC Processor is proposed to perform either prime field GF(p) operations or binary field GF(2m) operations for arbitrary prime numbers and irreducible polynomials. Besides, an application specific instruction set for ECC is presented to support parallel processing with VLIW instruction structure features and vector register addressing modes. Moreover, a new randomized clock cycle's insertion technique in regular calculation is designed against SPA and DPA attacks with 3% area and 4% power overhead. After implemented in 65-nm CMOS process, our proposed 521-bit dual field elliptic curve cryptographic Processor can perform scalar multiplication in 1.3 ms over GF(p521) and 0.94 ms over GF(2521). Our ECC Processor chip is advantageous not only in terms of functionality, scalability, and performance but also in protection against power-analysis attacks.
-
A High-Throughput Processor For Dual-Field Elliptic Curve Cryptography
2015 International Conference on Information and Communications Technologies (ICT 2015), 2015Co-Authors: Chaoma, Xiaohui Yang, Zibin DaiAbstract:In this paper, a mixture of VLIW and vector architecture of ECC Processor is proposed to perform either prime field GF(p) operations or binary field GF(2m) operations for arbitrary prime numbers and irreducible polynomials. Besides, an application specific instruction set for ECC is presented to support parallel processing with VLIW instruction structure features and vector register addressing modes. After implemented in 65-nm CMOS process, our proposed 521-bit dual field elliptic curve cryptographic Processor can perform scalar multiplication in 1.3 ms over GF(p521) and 0.94 ms over GF(2521). Our ECC Processor chip is advantageous in terms of functionality, scalability, and performance.
Mark Gebhart - One of the best experts on this subject based on the ideXlab platform.
-
unifying primary cache scratch and register file memories in a Throughput Processor
International Symposium on Microarchitecture, 2012Co-Authors: Mark Gebhart, Stephen W. Keckler, Brucek Khailany, Ronny Krashinsky, William J. DallyAbstract:Modern Throughput Processors such as GPUs employ thousands of threads to drive high-bandwidth, long-latency memory systems. These threads require substantial on-chip storage for registers, cache, and scratchpad memory. Existing designs hard-partition this local storage, fixing the capacities of these structures at design time. We evaluate modern GPU workloads and find that they have widely varying capacity needs across these different functions. Therefore, we propose a unified local memory which can dynamically change the partitioning among registers, cache, and scratchpad on a per-application basis. The tuning that this flexibility enables improves both performance and energy consumption, and broadens the scope of applications that can be efficiently executed on GPUs. Compared to a hard-partitioned design, we show that unified local memory provides a performance benefit as high as 71% along with an energy reduction up to 33%.
-
A Compile-Time Managed Multi-Level Register File Hierarchy
2012Co-Authors: Mark Gebhart, Stephen W. Keckler, William J. DallyAbstract:As Processors increasingly become power limited, performance improvements will be achieved by rearchitecting systems with energy efficiency as the primary design constraint. While some of these optimizations will be hardware based, combined hardware and software techniques likely will be the most productive. This work redesigns the register file system of a modern Throughput Processor with a combined hardware and software solution that reduces register file energy without harming system performance. Throughput Processors utilize a large number of threads to tolerate latency, requiring a large, energy-intensive register file to store thread context. Our results show that a compiler controlled register file hierarchy can reduce register file energy by up to 54%, compared to a hardware only caching approach that reduces register file energy by 34%. We explore register allocation algorithms that are specifically targeted to improve energy efficiency by sharing temporary register file resources across concurrently running threads and conduct a detailed limit study on the further potential to optimize operand delivery for Throughput Processors. Our efficiency gains represent a direct performance gain for power limited systems, such as GPUs
Stephen W. Keckler - One of the best experts on this subject based on the ideXlab platform.
-
unifying primary cache scratch and register file memories in a Throughput Processor
International Symposium on Microarchitecture, 2012Co-Authors: Mark Gebhart, Stephen W. Keckler, Brucek Khailany, Ronny Krashinsky, William J. DallyAbstract:Modern Throughput Processors such as GPUs employ thousands of threads to drive high-bandwidth, long-latency memory systems. These threads require substantial on-chip storage for registers, cache, and scratchpad memory. Existing designs hard-partition this local storage, fixing the capacities of these structures at design time. We evaluate modern GPU workloads and find that they have widely varying capacity needs across these different functions. Therefore, we propose a unified local memory which can dynamically change the partitioning among registers, cache, and scratchpad on a per-application basis. The tuning that this flexibility enables improves both performance and energy consumption, and broadens the scope of applications that can be efficiently executed on GPUs. Compared to a hard-partitioned design, we show that unified local memory provides a performance benefit as high as 71% along with an energy reduction up to 33%.
-
A Compile-Time Managed Multi-Level Register File Hierarchy
2012Co-Authors: Mark Gebhart, Stephen W. Keckler, William J. DallyAbstract:As Processors increasingly become power limited, performance improvements will be achieved by rearchitecting systems with energy efficiency as the primary design constraint. While some of these optimizations will be hardware based, combined hardware and software techniques likely will be the most productive. This work redesigns the register file system of a modern Throughput Processor with a combined hardware and software solution that reduces register file energy without harming system performance. Throughput Processors utilize a large number of threads to tolerate latency, requiring a large, energy-intensive register file to store thread context. Our results show that a compiler controlled register file hierarchy can reduce register file energy by up to 54%, compared to a hardware only caching approach that reduces register file energy by 34%. We explore register allocation algorithms that are specifically targeted to improve energy efficiency by sharing temporary register file resources across concurrently running threads and conduct a detailed limit study on the further potential to optimize operand delivery for Throughput Processors. Our efficiency gains represent a direct performance gain for power limited systems, such as GPUs
Chaoma - One of the best experts on this subject based on the ideXlab platform.
-
A High-Throughput Processor For Dual-Field Elliptic Curve Cryptography
2015 International Conference on Information and Communications Technologies (ICT 2015), 2015Co-Authors: Chaoma, Xiaohui Yang, Zibin DaiAbstract:In this paper, a mixture of VLIW and vector architecture of ECC Processor is proposed to perform either prime field GF(p) operations or binary field GF(2m) operations for arbitrary prime numbers and irreducible polynomials. Besides, an application specific instruction set for ECC is presented to support parallel processing with VLIW instruction structure features and vector register addressing modes. After implemented in 65-nm CMOS process, our proposed 521-bit dual field elliptic curve cryptographic Processor can perform scalar multiplication in 1.3 ms over GF(p521) and 0.94 ms over GF(2521). Our ECC Processor chip is advantageous in terms of functionality, scalability, and performance.