The Experts below are selected from a list of 17730 Experts worldwide ranked by ideXlab platform
Jeremy Kepner - One of the best experts on this subject based on the ideXlab platform.
-
PVTOL: Providing Productivity, Performance and Portability to DoD Signal Processing Applications on Multicore Processors
2008 DoD HPCMP Users Group Conference, 2008Co-Authors: Hahn Kim, Jeremy Kepner, Ryan Haney, Edward Rutledge, S. Sacco, Sanjeev Mohindra, M. Marzilli, J. Daly, Nadya T. BlissAbstract:PVTOL provides an object-oriented C++ API that hides the complexity of multicore architectures within a PGAS programming model, improving programmer productivity. Tasks and conduits enable data flow patterns such as pipelining and round-robining. Hierarchical maps concisely describe how to allocate hierarchical arrays across processor and memory hierarchies and provide a simple API for moving data across these hierarchies. Functors encapsulate computational kernels; new functors can be easily developed using the PVTOL API and can be fused for more efficient computation. Existing computation and communication technologies that are optimized for various architectures are used to achieve high performance. PVTOL abstracts the details of the underlying processor architectures to provide portability. We are actively developing PVTOL for Intel, PowerPC and Cell architectures and intend to add support for more computational kernels on these architectures. FPGAs are becoming popular for accelerating computation in both the high performance computing (HPC) and high performance embedded computing (HPEC) communities. Integrated processor-FPGA technologies are now available from both HPC and HPEC vendors, e.g. Cray and Mercury Computer Systems. We plan to support FPGAs as co-processors in PVTOL. Finally, automated mapping technology has been demonstrated with pMatlab. We plan to begin implementing automated mapping in PVTOL next year. Similar to PVL, as PVTOL matures and is used in more projects at Lincoln, we plan to propose concepts demonstrated in PVTOL to HPEC-SI for adoption into future versions of VSIPL++.
-
The HPEC Challenge Benchmark Suite
2006Co-Authors: Ryan Haney, Jeremy KepnerAbstract:Quantitative evaluation of different multi-processor High Performance embedded computing (HPEC) systems is an ongoing challenge for the HPEC community. The DARPA Polymorphous Computer Architecture (PCA) and HighProductivity computing Systems (HPCS) programs have created kernel and system level benchmarks and metrics for comparing the different architectures being developed under these programs. In this talk, we will describe a new benchmark suite drawn from the HPCS and PCA programs: the HPEC Challenge Benchmarks. It consists of eight single-processor kernel benchmarks and of a multiprocessor scalable synthetic SAR benchmark. We describe an implementation of the kernel benchmarks on the PowerPC G4 and the metrics used to evaluate it. We also demonstrate the parallel SAR benchmark and its scaling to multiple problem and machine sizes. The HPEC Challenge suite will be made widely available to community and will enable more rigorous comparison of HPEC systems.
-
parallel vsipl an open standard software library for high performance parallel signal processing
Proceedings of the IEEE, 2005Co-Authors: J M Lebak, Jeremy Kepner, Henry Hoffmann, Edward RutledgeAbstract:Real-time signal processing consumes the majority of the world's computing power. Increasingly, programmable parallel processors are used to address a wide variety of signal processing applications (e.g., scientific, video, wireless, medical, communication, encoding, radar, sonar, and imaging). In programmable systems, the major challenge is no longer hardware but software. Specifically, the key technical hurdle lies in allowing the user to write programs at high level, while still achieving performance and preserving the portability of the code across parallel computing hardware platforms. The Parallel Vector, Signal, and Image Processing Library (Parallel VSIPL++) addresses this hurdle by providing high-level C++ array constructs, a simple mechanism for mapping data and functions onto parallel hardware, and a community-defined portable interface. This paper presents an overview of the Parallel VSIPL++ standard as well as a deeper description of the technical foundations and expected performance of the library. Parallel VSIPL++ supports adaptive optimization at many levels. The C++ arrays are designed to support automatic hardware specialization by the compiler. The computation objects (e.g., fast Fourier transforms) are built with explicit setup and run stages to allow for runtime optimization. Parallel arrays and functions in Parallel VSIPL++ also support explicit setup and run stages, which are used to accelerate communication operations. The parallel mapping mechanism provides an external interface that allows optimal mappings to be generated offline and read into the system at runtime. Finally, the standard has been developed in collaboration with high performance embedded computing vendors and is compatible with their proprietary approaches to achieving performance.
-
Implementation of a Shipboard Ballistic Missile Defense Processing Application Using the High Performance embedded computing Software Initiative (HPEC-SI) API
2004Co-Authors: Joseph Cook, Nathan Doss, Jane Kent, Rick Pancoast, Jordan Lusterman, Jeremy KepnerAbstract:Abstract : This briefing describes an effort to implement advanced Shipboard Ballistic Missile Defense (SBMD) application algorithms utilizing HPEC-SI. Shipboard application code, previously written in the C programming language for conventional COTS PowerPC-based embedded architectures, is being converted by Lockheed Martin MS2, as an HPEC-SI Demonstration, to run under the HPEC-SI API. The C code, designed to run in a C environment, will be converted to the HPEC-SI API standard to run under a true C++ Object Oriented environment, and will eventually take advantage of the HPEC-SI parallel processing features. Of particular interest in this conversion is a comparison of key DoD processing algorithms executed on a conventional, embedded processing architecture using C and C application libraries, as compared with execution in an embedded HPEC-SI processing environment. In this briefing, we describe the porting of several of the signal processing algorithms that have been developed using C-based VSIPL, and port them to the HPEC-SI VSIPL++ API under development. As part of this process, the HPEC-SI community will receive valuable feedback regarding the HPEC-SI API implementation, including the development process, development metrics, development environment issues and key library functions. Eventually, the open architecture HPEC-SI VSIPL++ code developed for the Navy and MDA will be ported to a tactical system for deployment on Aegis cruisers and destroyers.
-
High Performance embedded computing Software Initiative (HPEC-SI)
2004Co-Authors: Jeremy KepnerAbstract:Abstract : The High Performance embedded computing Software Initiative is addressing the military's need to advance the state of embedded software development tools, libraries, and methodologies to retain the nation's military technology advantage in increasingly software-based systems. Key accomplishments include the completion of the first demonstration and the development of the Parallel VSIPL++ standard. Currently, the HPEC-SI effort is on track towards its goal of changing the state-of-the-practice in programming DoD HPEC SIP systems. This paper gives a brief overview of the HPEC-SI program objectives, technical objectives and program plans. The HPEC-SI program is organized around demonstrations, standards development, and applied research. Each of these activities is overseen by a Working Group. The demonstrations team Prime contractors with FFRDC or academic partners to use currently defined standards, evaluate their performance, and report on how well their needs are being met. The first demonstration was with the Common Imagery Processor (CIP) and successfully showed the use of MPI communication standard and the VSIPL computation standard to achieve portability (while preserving performance) across shared servers and distributed memory embedded systems. The Development Working Group is extending the VSIPL standard to include parallel object-oriented software practices already prototyped by the research community. This effort is tightly coupled with military demonstrations; it provides the next generation of standards with direct feedback from the military user base. The Applied Research Working Group is also taking a longer term view to assess the potential impact of a variety of emerging technologies, such as fault tolerance and dynamic scheduling, self-optimization, and next generation, high productivity languages. Thirty-one briefing charts summarize the presentation.
Hoai Hoang - One of the best experts on this subject based on the ideXlab platform.
-
IPDPS - A fibre-optic AWG-based real-time network and its applicability to high-performance embedded computing
19th IEEE International Parallel and Distributed Processing Symposium, 1Co-Authors: Annette Böhm, Magnus Jonsson, Kristina Kunert, Hoai HoangAbstract:In this paper, an architecture and a medium access control (MAC) protocol for a multi-wavelength optical communication network, applicable in short range communication systems like system area networks (SANs), are proposed. The main focus lies on guaranteed support for hard and soft real-time traffic. The network is based upon a single-hop star topology with an arrayed waveguide grating (AWG) at its center. Traffic scheduling is centralized in one node (residing together with the AWG in a hub), which communicates through a physical control channel. The AWG's property of spatial wavelength reuse and the combination of fixed-tuned and tunable transceivers in the nodes enable simultaneous control and data transmission. A case study with defined real-time communication requirements in the field of radar signal processing (RSP) was carried out and indicates that the proposed system is very suitable for this kind of application.
Edward Rutledge - One of the best experts on this subject based on the ideXlab platform.
-
PVTOL: Providing Productivity, Performance and Portability to DoD Signal Processing Applications on Multicore Processors
2008 DoD HPCMP Users Group Conference, 2008Co-Authors: Hahn Kim, Jeremy Kepner, Ryan Haney, Edward Rutledge, S. Sacco, Sanjeev Mohindra, M. Marzilli, J. Daly, Nadya T. BlissAbstract:PVTOL provides an object-oriented C++ API that hides the complexity of multicore architectures within a PGAS programming model, improving programmer productivity. Tasks and conduits enable data flow patterns such as pipelining and round-robining. Hierarchical maps concisely describe how to allocate hierarchical arrays across processor and memory hierarchies and provide a simple API for moving data across these hierarchies. Functors encapsulate computational kernels; new functors can be easily developed using the PVTOL API and can be fused for more efficient computation. Existing computation and communication technologies that are optimized for various architectures are used to achieve high performance. PVTOL abstracts the details of the underlying processor architectures to provide portability. We are actively developing PVTOL for Intel, PowerPC and Cell architectures and intend to add support for more computational kernels on these architectures. FPGAs are becoming popular for accelerating computation in both the high performance computing (HPC) and high performance embedded computing (HPEC) communities. Integrated processor-FPGA technologies are now available from both HPC and HPEC vendors, e.g. Cray and Mercury Computer Systems. We plan to support FPGAs as co-processors in PVTOL. Finally, automated mapping technology has been demonstrated with pMatlab. We plan to begin implementing automated mapping in PVTOL next year. Similar to PVL, as PVTOL matures and is used in more projects at Lincoln, we plan to propose concepts demonstrated in PVTOL to HPEC-SI for adoption into future versions of VSIPL++.
-
parallel vsipl an open standard software library for high performance parallel signal processing
Proceedings of the IEEE, 2005Co-Authors: J M Lebak, Jeremy Kepner, Henry Hoffmann, Edward RutledgeAbstract:Real-time signal processing consumes the majority of the world's computing power. Increasingly, programmable parallel processors are used to address a wide variety of signal processing applications (e.g., scientific, video, wireless, medical, communication, encoding, radar, sonar, and imaging). In programmable systems, the major challenge is no longer hardware but software. Specifically, the key technical hurdle lies in allowing the user to write programs at high level, while still achieving performance and preserving the portability of the code across parallel computing hardware platforms. The Parallel Vector, Signal, and Image Processing Library (Parallel VSIPL++) addresses this hurdle by providing high-level C++ array constructs, a simple mechanism for mapping data and functions onto parallel hardware, and a community-defined portable interface. This paper presents an overview of the Parallel VSIPL++ standard as well as a deeper description of the technical foundations and expected performance of the library. Parallel VSIPL++ supports adaptive optimization at many levels. The C++ arrays are designed to support automatic hardware specialization by the compiler. The computation objects (e.g., fast Fourier transforms) are built with explicit setup and run stages to allow for runtime optimization. Parallel arrays and functions in Parallel VSIPL++ also support explicit setup and run stages, which are used to accelerate communication operations. The parallel mapping mechanism provides an external interface that allows optimal mappings to be generated offline and read into the system at runtime. Finally, the standard has been developed in collaboration with high performance embedded computing vendors and is compatible with their proprietary approaches to achieving performance.
-
an open standard software library for high performance parallel signal processing the parallel vsipl library
2004Co-Authors: J M Lebak, Jeremy Kepner, Henry Hoffmann, Edward RutledgeAbstract:Real-time signal processing consumes the majority of the world’s computing power. Increasingly, programmable parallel processors are used to address a wide variety of signal processing applications (e.g. scientific, video, wireless, medical, communication, encoding, radar, sonar and imaging). In programmable systems the major challenge is no longer hardware but software. Specifically, the key technical hurdle lies in allowing the user to write programs at high level, while still achieving performance and preserving the portability of the code across parallel computing hardware platforms. The Parallel Vector, Signal, and Image Processing Library (Parallel VSIPL++) addresses this hurdle by providing high level C++ array constructs, a simple mechanism for mapping data and functions onto parallel hardware, and a community-defined portable interface. This paper presents an overview of the Parallel VSIPL++ standard as well as a deeper description of the technical foundations and expected performance of the library. Parallel VSIPL++ supports adaptive optimization at many levels. The C++ arrays are designed to support automatic hardware specialization by the compiler. The computation objects (e.g. Fast Fourier Transforms) are built with explicit setup and run stages to allow for run-time optimization. Parallel arrays and functions in VSIPL++ support these same features, which are used to accelerate communication operations. The parallel mapping mechanism provides an external interface that allows optimal mappings to be generated off-line and read into the system at run-time. Finally, the standard has been developed in collaboration with high performance embedded computing vendors and is compatible with their proprietary approaches to achieving performance. Copyright 2004 MIT Lincoln Laboratory. This work is sponsored by the High Performance computing Modernization Office under Air Force Contract F19628-00-C-0002. Opinions, interpretations, conclusions, and recommendations are those of the authors and are not necessarily endorsed by the United States Government. April 15, 2004 DRAFT
Marco Platzner - One of the best experts on this subject based on the ideXlab platform.
-
a hardware software infrastructure for performance monitoring on leon3 multicore platforms
Field-Programmable Logic and Applications, 2014Co-Authors: Nam Ho, Paul Kaufmann, Marco PlatznerAbstract:Monitoring applications at run-time and evaluating the recorded statistical data of the underlying micro architecture is one of the key aspects required by many hardware architects and system designers as well as high-performance software developers. To fulfill this requirement, most modern CPUs for High Performance computing have been equipped with Performance Monitoring Units (PMU) including a set of hardware counters, which can be configured to monitor a rich set of events. Unfortunately, embedded and reconfigurable systems are mostly lacking this feature. Towards rapid exploration of High Performance embedded computing in near future, we believe that supporting PMU for these systems is necessary. In this paper, we propose a PMU infrastructure, which supports monitoring of up to seven concurrent events. The PMU infrastructure is implemented on an FPGA and is integrated into a LEON3 platform.We show also the integration of our PMU infrastructure with the perf_event, which is the standard PMU architecture of the Linux kernel.
-
FPL - A hardware/software infrastructure for performance monitoring on LEON3 multicore platforms
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Paul Kaufmann, Marco PlatznerAbstract:Monitoring applications at run-time and evaluating the recorded statistical data of the underlying micro architecture is one of the key aspects required by many hardware architects and system designers as well as high-performance software developers. To fulfill this requirement, most modern CPUs for High Performance computing have been equipped with Performance Monitoring Units (PMU) including a set of hardware counters, which can be configured to monitor a rich set of events. Unfortunately, embedded and reconfigurable systems are mostly lacking this feature. Towards rapid exploration of High Performance embedded computing in near future, we believe that supporting PMU for these systems is necessary. In this paper, we propose a PMU infrastructure, which supports monitoring of up to seven concurrent events. The PMU infrastructure is implemented on an FPGA and is integrated into a LEON3 platform.We show also the integration of our PMU infrastructure with the perf_event, which is the standard PMU architecture of the Linux kernel.
Annette Böhm - One of the best experts on this subject based on the ideXlab platform.
-
A fibre-optic AWG-based real-time network for high-performance embedded computing
2004Co-Authors: Annette Böhm, Magnus Jonsson, Kristina KunertAbstract:In this paper, an architecture and a Medium Access Control (MAC) protocol for a multiwavelength optical communication network, applicable in short range communication systems like System Area Netwo ...
-
IPDPS - A fibre-optic AWG-based real-time network and its applicability to high-performance embedded computing
19th IEEE International Parallel and Distributed Processing Symposium, 1Co-Authors: Annette Böhm, Magnus Jonsson, Kristina Kunert, Hoai HoangAbstract:In this paper, an architecture and a medium access control (MAC) protocol for a multi-wavelength optical communication network, applicable in short range communication systems like system area networks (SANs), are proposed. The main focus lies on guaranteed support for hard and soft real-time traffic. The network is based upon a single-hop star topology with an arrayed waveguide grating (AWG) at its center. Traffic scheduling is centralized in one node (residing together with the AWG in a hub), which communicates through a physical control channel. The AWG's property of spatial wavelength reuse and the combination of fixed-tuned and tunable transceivers in the nodes enable simultaneous control and data transmission. A case study with defined real-time communication requirements in the field of radar signal processing (RSP) was carried out and indicates that the proposed system is very suitable for this kind of application.