The Experts below are selected from a list of 7140 Experts worldwide ranked by ideXlab platform
Josep Torrellas - One of the best experts on this subject based on the ideXlab platform.
-
The Need for Fast Communication in Hardware-Based Speculative Chip Multiprocessors
International Journal of Parallel Programming, 2001Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Chip-multiprocessor (CMP) architectures are a promising design alternative to exploit the ever-increasing number of transistors that can be put on a die. To deliver high performance on applications that cannot be easily parallelized, CMPs can use additional support for speculatively executing the possibly data-dependent threads of an application. For cross-thread dependences that must be handled dynamically, the threads can be made to synchronize and communicate either at the Register level or at the memory level. In the past, it has been unclear whether the higher hardware cost of Register-level communication is cost-effective. In this paper, we show that the wide-issue dynamic processors that will soon populate CMPs, make fast communication a requirement for high performance. Consequently, we propose an effective hardware mechanism to support communication and synchronization of Registers between on-chip processors. Our scheme adds enough support to Enable Register-level communication without specializing the architecture toward speculation much. Finally, our scheme allows the system to achieve near ideal performance.
-
A Chip-Multiprocessor Architecture with
1999Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Much emphasis is now placed on chip-multiprocessor (CMP) architectures for exploiting thread-level parallelism in an application. In such architectures, speculation may be employed to execute applications that cannot be parallelized statically. In this paper, we present an efficient CMP architecture for speculative execution of sequential binaries without source recompilation. We present the software support that Enables identification of threads from a sequential binary. The hardware includes a memory disambiguation mechanism that Enables the detection of interthread memory dependence violations during speculative execution. This hardware is different from past proposals in that it does not rely on a snoopy-based cache-coherence protocol. Instead, it uses an approach similar to a directory-based scheme. Furthermore, the architecture includes a simple and efficient hardware mechanism to Enable Register-level communication between on-chip processors. Evaluation of this software-hardware approach shows that it is quite effective in achieving high performance when running sequential binaries.
-
A chip-multiprocessor architecture with speculative multithreading
IEEE Transactions on Computers, 1999Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Much emphasis is now being placed on chip-multiprocessor (CMP) architectures for exploiting thread-level parallelism in applications. In such architectures, speculation may be employed to execute applications that cannot be parallelized statically. In this paper, we present an efficient CMP architecture for the speculative execution of sequential binaries without source recompilation. We present software support that Enables the identification of threads from a sequential binary. The hardware includes a memory disambiguation mechanism that Enables the detection of interthread memory dependence violations during speculative execution. This hardware is different from past proposals in that it does not rely on a snoopy-based cache-coherence protocol. Instead, it uses an approach similar to a directory-based scheme. Furthermore, the architecture includes a simple and efficient hardware mechanism to Enable Register-level communication between on-chip processors. Evaluation of this software-hardware approach shows that it is quite effective in achieving high performance when running sequential binaries.
-
IEEE PACT - The need for fast communication in hardware-based speculative chip multiprocessors
1999 International Conference on Parallel Architectures and Compilation Techniques (Cat. No.PR00425), 1Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Chip-multiprocessor (CMP) architectures are a promising design alternative to exploit the ever-increasing number of transistors that can be put on a die. To deliver high performance on applications that cannot be easily parallelized, CMPs can use additional support for speculatively executing the possibly data-dependent threads of an application. While some of the cross-thread dependences in applications must be handled dynamically, others can be fully determined by the compiler. For the latter dependences, the threads can be made to synchronize and communicate either at the Register level or at the memory level. In the past, it has been unclear whether the higher hardware cost of Register-level communication is cost-effective. In this paper, we show that the wide-issue dynamic processors that will soon populate CMPs, make fast communication a requirement for high performance. Consequently, we propose an effective hardware mechanism to support communication and synchronization of Registers between on-chip processors. Our scheme adds enough support to Enable Register-level communication without specializing the architecture so much toward speculation that it leads to much unutilized hardware under workloads that do not need speculative parallelization. Finally, the scheme allows the system to achieve near ideal performance.
Venkata Krishnan - One of the best experts on this subject based on the ideXlab platform.
-
The Need for Fast Communication in Hardware-Based Speculative Chip Multiprocessors
International Journal of Parallel Programming, 2001Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Chip-multiprocessor (CMP) architectures are a promising design alternative to exploit the ever-increasing number of transistors that can be put on a die. To deliver high performance on applications that cannot be easily parallelized, CMPs can use additional support for speculatively executing the possibly data-dependent threads of an application. For cross-thread dependences that must be handled dynamically, the threads can be made to synchronize and communicate either at the Register level or at the memory level. In the past, it has been unclear whether the higher hardware cost of Register-level communication is cost-effective. In this paper, we show that the wide-issue dynamic processors that will soon populate CMPs, make fast communication a requirement for high performance. Consequently, we propose an effective hardware mechanism to support communication and synchronization of Registers between on-chip processors. Our scheme adds enough support to Enable Register-level communication without specializing the architecture toward speculation much. Finally, our scheme allows the system to achieve near ideal performance.
-
A Chip-Multiprocessor Architecture with
1999Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Much emphasis is now placed on chip-multiprocessor (CMP) architectures for exploiting thread-level parallelism in an application. In such architectures, speculation may be employed to execute applications that cannot be parallelized statically. In this paper, we present an efficient CMP architecture for speculative execution of sequential binaries without source recompilation. We present the software support that Enables identification of threads from a sequential binary. The hardware includes a memory disambiguation mechanism that Enables the detection of interthread memory dependence violations during speculative execution. This hardware is different from past proposals in that it does not rely on a snoopy-based cache-coherence protocol. Instead, it uses an approach similar to a directory-based scheme. Furthermore, the architecture includes a simple and efficient hardware mechanism to Enable Register-level communication between on-chip processors. Evaluation of this software-hardware approach shows that it is quite effective in achieving high performance when running sequential binaries.
-
A chip-multiprocessor architecture with speculative multithreading
IEEE Transactions on Computers, 1999Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Much emphasis is now being placed on chip-multiprocessor (CMP) architectures for exploiting thread-level parallelism in applications. In such architectures, speculation may be employed to execute applications that cannot be parallelized statically. In this paper, we present an efficient CMP architecture for the speculative execution of sequential binaries without source recompilation. We present software support that Enables the identification of threads from a sequential binary. The hardware includes a memory disambiguation mechanism that Enables the detection of interthread memory dependence violations during speculative execution. This hardware is different from past proposals in that it does not rely on a snoopy-based cache-coherence protocol. Instead, it uses an approach similar to a directory-based scheme. Furthermore, the architecture includes a simple and efficient hardware mechanism to Enable Register-level communication between on-chip processors. Evaluation of this software-hardware approach shows that it is quite effective in achieving high performance when running sequential binaries.
-
IEEE PACT - The need for fast communication in hardware-based speculative chip multiprocessors
1999 International Conference on Parallel Architectures and Compilation Techniques (Cat. No.PR00425), 1Co-Authors: Venkata Krishnan, Josep TorrellasAbstract:Chip-multiprocessor (CMP) architectures are a promising design alternative to exploit the ever-increasing number of transistors that can be put on a die. To deliver high performance on applications that cannot be easily parallelized, CMPs can use additional support for speculatively executing the possibly data-dependent threads of an application. While some of the cross-thread dependences in applications must be handled dynamically, others can be fully determined by the compiler. For the latter dependences, the threads can be made to synchronize and communicate either at the Register level or at the memory level. In the past, it has been unclear whether the higher hardware cost of Register-level communication is cost-effective. In this paper, we show that the wide-issue dynamic processors that will soon populate CMPs, make fast communication a requirement for high performance. Consequently, we propose an effective hardware mechanism to support communication and synchronization of Registers between on-chip processors. Our scheme adds enough support to Enable Register-level communication without specializing the architecture so much toward speculation that it leads to much unutilized hardware under workloads that do not need speculative parallelization. Finally, the scheme allows the system to achieve near ideal performance.
Kevin O'brien - One of the best experts on this subject based on the ideXlab platform.
-
SC - Compact multi-dimensional kernel extraction for Register tiling
Proceedings of the Conference on High Performance Computing Networking Storage and Analysis - SC '09, 2009Co-Authors: Lakshminarayanan Renganarayana, Uday Bondhugula, Salem Derisavi, Alexandre E. Eichenberger, Kevin O'brienAbstract:To achieve high performance on multi-cores, modern loop optimizers apply long sequences of transformations that produce complex loop structures. Downstream optimizations such as Register tiling (unroll-and-jam plus scalar promotion) typically provide a significant performance improvement. Typical Register tilers provide this performance improvement only when applied on simple loop structures. They often fail to operate on complex loop structures leaving a significant amount of performance on the table. We present a technique called compact multi-dimensional kernel extraction (COMDEX) which can make Register tilers operate on arbitrarily complex loop structures and Enable them to provide the performance benefits. COMDEX extracts compact unrollable kernels from complex loops. We show that by using COMDEX as a pre-processing to Register tiling we can (i) Enable Register tiling on complex loop structures and (ii) realize a significant performance improvement on a variety of codes.
Kaushik Roy - One of the best experts on this subject based on the ideXlab platform.
-
A 200mV to 1.2V, 4.4MHz to 6.3GHz, 48×42b 1R/1W programmable Register file in 65nm CMOS
ESSCIRC 2007 - 33rd European Solid-State Circuits Conference, 2007Co-Authors: Amit Agarwal, N. Banerjee, Steven K. Hsu, Ram Krishnamurthy, Kaushik RoyAbstract:This paper describes a 48times42b 1-read, 1-write ported Register file which operates at supply voltage range of 1.2 V (6.1-6.3 GHz, 47 mW) down to 0.2 V (4-4.4 MHz, 0.01 mW) in 65 nm CMOS. Two programmable techniques, triple stacking and forced stacking are proposed which Enable Register files to operate at ultra low supply voltages while maintaining the performance comparable to conventional design at high supply voltages.
Brian J. D’auriol - One of the best experts on this subject based on the ideXlab platform.
-
All-optical Linear Array with a Reconfigurable Pipelined Bus System (OLARPBS) optical bus parallel computing model
The Journal of Supercomputing, 2016Co-Authors: Brian J. D’auriolAbstract:The All-optical Linear Array with a Reconfigurable Pipelined Bus System (OLARPBS) optical bus parallel computing model is proposed in this paper. The OLARPBS model includes several architectural and logical extensions to the existing LARPBS(p) model. Architecturally, the extensions include the replacement of the electronic processors by fine-grained, all-optical digital processing elements; and the replacement of the optical conduit (optical bus) by one or more optical conduits, each consisting of a structural hierarchy of bundles of light paths that Enable multiple parallel and/or pipelined message communications. Logically, the extensions include: parallel and/or pipelined, conduit-based, optical Registers; Register scheduling; and processor data path requirements. Collectively, these extensions Enable Register-based bits-in-flight algorithm development. A study of two applications with implementation on the OLARPBS model is considered.