The Experts below are selected from a list of 16644 Experts worldwide ranked by ideXlab platform
David A Richie - One of the best experts on this subject based on the ideXlab platform.
-
parallel programming model for the epiphany many core coProcessor using threaded mpi
Microprocessors and Microsystems, 2016Co-Authors: James A Ross, David A Richie, Song Jun Park, Dale R ShiresAbstract:We investigate the use of MPI for programming the Epiphany RISC Array Processor.A threaded MPI implementation adapted for coProcessor offload is presented.Existing MPI code for four scientific applications was re-used with minimal changes.Demonstrated performance exceeds 12 GFLOPS with an efficiency over 20GFLOPS/W.Threaded MPI exhibits the highest performance reported using a standard parallel API. The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC Array Processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.
-
implementing openshmem for the adapteva epiphany risc Array Processor
International Conference on Conceptual Structures, 2016Co-Authors: James A Ross, David A RichieAbstract:The energy-efficient Adapteva Epiphany architecture exhibits massive many-core scalability in a physically compact 2D Array of RISC cores with a fast network-on-chip (NoC). With fully divergent cores capable of MIMD execution, the physical topology and memory-mapped capabilities of the core and network translate well to partitioned global address space (PGAS) parallel programming models. Following an investigation into the use of two-sided communication using threaded MPI, one-sided communication using SHMEM is being explored. Here we present work in progress on the development of an OpenSHMEM 1.2 implementation for the Epiphany architecture.
-
threaded mpi programming model for the epiphany risc Array Processor
Journal of Computational Science, 2015Co-Authors: David A Richie, James A Ross, Song Jun Park, Dale R ShiresAbstract:The low-power Adapteva Epiphany RISC Array Processor offers high computational energy-efficiency and parallel scalability. However, extracting performance with a standard parallel programming model remains a great challenge. We present an effective programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard adapted for coProcessor offload. Using MPI exploits the similarities between the Epiphany architecture and a networked parallel distributed cluster. Furthermore, our approach enables codes written with MPI to execute on the RISC Array Processor with little modification. We present experimental results for matrix–matrix multiplication using MPI and highlight the importance of fast inter-core data transfers. Using MPI we demonstrate an on-chip performance of 9.1 GFLOPS with an efficiency of 15.3 GFLOPS/W. Threaded MPI exhibits the highest performance reported for the Epiphany architecture using a standard parallel programming model.
-
parallel programming model for the epiphany many core coProcessor using threaded mpi
arXiv: Distributed Parallel and Cluster Computing, 2015Co-Authors: James A Ross, David A Richie, Song Jun Park, Dale R ShiresAbstract:The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC Array Processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, a five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.
P A Ivey - One of the best experts on this subject based on the ideXlab platform.
-
an Array Processor for general purpose digital image compression
IEEE Journal of Solid-state Circuits, 1995Co-Authors: R B Yates, Neil A Thacker, S J Evans, S N Walker, P A IveyAbstract:A new VLSI Processor (DIP chip) for image compression is presented which combines principles of multipipeline and Array processing. The device is not specific to any one image compression algorithm and can be regarded as a general purpose Processor. The chip has been implemented using a CMOS 1.0-/spl mu/m process on a 14.4/spl times/13.5-mm/sup 2/ die. An internal clock frequency of 40 MHz results in 1.2/spl times/10/sup 9/ operations/s on 8-bit data. Solutions to problems associated with the large bandwidth required, for both image data and instruction streams, is the main aim of the paper. The necessary problem of increasing the Array clock frequency relative to the input/output clock frequency without the need for a large on-chip instruction cache or fast external clock speeds is also addressed. >
-
an Array Processor for general purpose digital image compression
Custom Integrated Circuits Conference, 1995Co-Authors: R B Yates, Neil A Thacker, S J Evans, S N Walker, P A IveyAbstract:―A new VLSI Processor (DIP chip) for image compression is presented which combines principles of multipipeline and Array processing. The device is not specific to any one image compression algorithm and can be regarded as a general purpose Processor. The chip has been implemented using a CMOS 1.0-μm process on a 14.4 x 13.5-mm 2 die. An internal clock frequency of 40 MHz results in 1.2 x 10 9 operations/s on 8-bit data. Solutions to problems associated with the large bandwidth required, for both image data and instruction streams, is the main aim of the paper. The necessary problem of increasing the Array clock frequency relative to the input/output clock frequency without the need for a large on-chip instruction cache or fast external clock speeds is also addressed.
Tamio Arai - One of the best experts on this subject based on the ideXlab platform.
-
an integrated memory Array Processor for embedded image recognition systems
IEEE Transactions on Computers, 2007Co-Authors: S. Okazaki, Tamio AraiAbstract:Embedded Processors for video image recognition in most cases not only need to address the conventional cost (die size and power) versus real-time performance issue, but must also maintain high flexibility due to the immense diversity of recognition targets, situations, and applications. This paper describes IMAP, a highly parallel SIMD linear Processor and memory Array architecture that addresses these trade-off requirements. By using parallel and systolic algorithmic techniques, but based on a simple linear Array architecture, IMAP successfully exploits not only the straightforward per-image row data level parallelism (DLP), but also the inherent DLP of other memory access patterns frequently found in various image recognition tasks, while allowing programming to be done using an explicit parallel C language (1DC). We describe and evaluate IMAP-CE, one of the latest IMAP Processors, integrating 128 100 MHz 8 bit 4-way VLIW PEs, 128 2 KByte RAMs, and one 16 bit RISC control Processor onto a single chip. The PE instruction set is enhanced to support 1DC code. The die size of IMAP-CE is 11 times11 mm2 integrating 32.7 M transistors, while the power consumption is, on average, approximately 2 watts. IMAP-CE is evaluated mainly by comparing its performance while running 1DC code with that of a 2.4 GHz Intel P4 running optimized C code. Based on the use of parallelizing techniques, benchmark results show a speed increase of up to 20 times for image filter kernels and of 4 times for a full image recognition application
-
an integrated memory Array Processor architecture for embedded image recognition systems
International Symposium on Computer Architecture, 2005Co-Authors: Shorin Kyo, Shinichiro Okazaki, Tamio AraiAbstract:Embedded Processors for video image recognition require to address both the cost (die size and power) versus real-time performance issue, and also to achieve high flexibility due to the immense diversity of recognition targets, situations, and applications. This paper describes IMAP, a highly parallel SIMD linear Processor and memory Array architecture that addresses these trading-off requirements. By using parallel and systolic algorithmic techniques, despite of its simple architecture IMAP achieves to exploit not only the straightforward per image row data level parallelism (DLP), but also the inherent DLP of other memory access patterns frequently found in various image recognition tasks, under the use of an explicit parallel C language (1DC). We describe and evaluate IMAP-CE, a latest IMAP Processor, which integrates 128 of 100MHz 8 bit4-way VLIW PEs, 128 of 2KByte RAMs, and one 16 bit RISC control Processor, into a single chip. The PE instruction set is enhanced for supporting 1DC codes. IMAP-CE is evaluated mainly by comparing its performance running 1DC codes with that of a 2.4GHz Intel P4 running optimized C codes. Based on the use of parallelizing techniques, benchmark results show a speedup of up to 20 for image filter kernels, and of 4 for a full image recognition application.
James A Ross - One of the best experts on this subject based on the ideXlab platform.
-
parallel programming model for the epiphany many core coProcessor using threaded mpi
Microprocessors and Microsystems, 2016Co-Authors: James A Ross, David A Richie, Song Jun Park, Dale R ShiresAbstract:We investigate the use of MPI for programming the Epiphany RISC Array Processor.A threaded MPI implementation adapted for coProcessor offload is presented.Existing MPI code for four scientific applications was re-used with minimal changes.Demonstrated performance exceeds 12 GFLOPS with an efficiency over 20GFLOPS/W.Threaded MPI exhibits the highest performance reported using a standard parallel API. The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC Array Processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.
-
implementing openshmem for the adapteva epiphany risc Array Processor
International Conference on Conceptual Structures, 2016Co-Authors: James A Ross, David A RichieAbstract:The energy-efficient Adapteva Epiphany architecture exhibits massive many-core scalability in a physically compact 2D Array of RISC cores with a fast network-on-chip (NoC). With fully divergent cores capable of MIMD execution, the physical topology and memory-mapped capabilities of the core and network translate well to partitioned global address space (PGAS) parallel programming models. Following an investigation into the use of two-sided communication using threaded MPI, one-sided communication using SHMEM is being explored. Here we present work in progress on the development of an OpenSHMEM 1.2 implementation for the Epiphany architecture.
-
threaded mpi programming model for the epiphany risc Array Processor
Journal of Computational Science, 2015Co-Authors: David A Richie, James A Ross, Song Jun Park, Dale R ShiresAbstract:The low-power Adapteva Epiphany RISC Array Processor offers high computational energy-efficiency and parallel scalability. However, extracting performance with a standard parallel programming model remains a great challenge. We present an effective programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard adapted for coProcessor offload. Using MPI exploits the similarities between the Epiphany architecture and a networked parallel distributed cluster. Furthermore, our approach enables codes written with MPI to execute on the RISC Array Processor with little modification. We present experimental results for matrix–matrix multiplication using MPI and highlight the importance of fast inter-core data transfers. Using MPI we demonstrate an on-chip performance of 9.1 GFLOPS with an efficiency of 15.3 GFLOPS/W. Threaded MPI exhibits the highest performance reported for the Epiphany architecture using a standard parallel programming model.
-
parallel programming model for the epiphany many core coProcessor using threaded mpi
arXiv: Distributed Parallel and Cluster Computing, 2015Co-Authors: James A Ross, David A Richie, Song Jun Park, Dale R ShiresAbstract:The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC Array Processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, a five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.
Dale R Shires - One of the best experts on this subject based on the ideXlab platform.
-
parallel programming model for the epiphany many core coProcessor using threaded mpi
Microprocessors and Microsystems, 2016Co-Authors: James A Ross, David A Richie, Song Jun Park, Dale R ShiresAbstract:We investigate the use of MPI for programming the Epiphany RISC Array Processor.A threaded MPI implementation adapted for coProcessor offload is presented.Existing MPI code for four scientific applications was re-used with minimal changes.Demonstrated performance exceeds 12 GFLOPS with an efficiency over 20GFLOPS/W.Threaded MPI exhibits the highest performance reported using a standard parallel API. The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC Array Processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.
-
threaded mpi programming model for the epiphany risc Array Processor
Journal of Computational Science, 2015Co-Authors: David A Richie, James A Ross, Song Jun Park, Dale R ShiresAbstract:The low-power Adapteva Epiphany RISC Array Processor offers high computational energy-efficiency and parallel scalability. However, extracting performance with a standard parallel programming model remains a great challenge. We present an effective programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard adapted for coProcessor offload. Using MPI exploits the similarities between the Epiphany architecture and a networked parallel distributed cluster. Furthermore, our approach enables codes written with MPI to execute on the RISC Array Processor with little modification. We present experimental results for matrix–matrix multiplication using MPI and highlight the importance of fast inter-core data transfers. Using MPI we demonstrate an on-chip performance of 9.1 GFLOPS with an efficiency of 15.3 GFLOPS/W. Threaded MPI exhibits the highest performance reported for the Epiphany architecture using a standard parallel programming model.
-
parallel programming model for the epiphany many core coProcessor using threaded mpi
arXiv: Distributed Parallel and Cluster Computing, 2015Co-Authors: James A Ross, David A Richie, Song Jun Park, Dale R ShiresAbstract:The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC Array Processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, a five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.