The Experts below are selected from a list of 22725 Experts worldwide ranked by ideXlab platform
Steven Swanson - One of the best experts on this subject based on the ideXlab platform.
-
onyx a protoype phase change memory Storage Array
USENIX conference on Hot topics in storage and file systems, 2011Co-Authors: Ameen D Akel, Adrian M Caulfield, Todor Mollov, Rajesh Gupta, Steven SwansonAbstract:We describe a prototype high-performance solid-state drive based on first-generation phase-change memory (PCM) devices called Onyx. Onyx has a capacity of 10 GB and connects to the host system via PCIe. We describe the internal architecture of Onyx including the PCM memory modules we constructed and the FPGA-based controller that manages them. Onyx can perform a 4 KB random read in 38 µs and sustain 191K 4 KB read IO operations per second. A 4 KB write requires 179 µs. We describe our experience tuning the Onyx system to reduce the cost of wear-leveling and increase performance. We find that Onyx out-performs a state-of-the-art flash-based SSD for small writes (< 2 KB) by between 72 and 120% and for reads of all sizes. In addition, Onyx incurs 20-51% less CPU overhead per IOP for small requests. Combined, our results demonstrate that even first-generation PCM SSDs can outperform flash-based Arrays for the irregular (and frequently read-dominated) access patterns that define many of today's "killer" Storage applications. Next generation PCM devices will widen the performance gap further and set the stage for PCM becoming a serious flash competitor in many applications.
-
moneta a high performance Storage Array architecture for next generation non volatile memories
International Symposium on Microarchitecture, 2010Co-Authors: Adrian M Caulfield, Rajesh Gupta, Joel Coburn, Todor I Mollow, Steven SwansonAbstract:Emerging non-volatile memory technologies such as phase change memory (PCM) promise to increase Storage system performance by a wide margin relative to both conventional disks and flash-based SSDs. Realizing this potential will require significant changes to the way systems interact with Storage devices as well as a rethinking of the Storage devices themselves. This paper describes the architecture of a prototype PCIe-attached Storage Array built from emulated PCM Storage called Moneta. Moneta provides a carefully designed hardware/software interface that makes issuing and completing accesses atomic. The atomic management interface, combined with hardware scheduling optimizations, and an optimized Storage stack increases performance for small, random accesses by 18x and reduces software overheads by 60%. Moneta Array sustain 2.8~GB/s for sequential transfers and 541K random 4~KB~IO operations per second (8x higher than a state-of-the-art flash-based SSD). Moneta can perform a 512-byte write in 9~us (5.6x faster than the SSD). Moneta provides a harmonic mean speedup of 2.1x and a maximum speed up of 9x across a range of file system, paging, and database workloads. We also explore trade-offs in Moneta's architecture between performance, power, memory organization, and memory latency.
Adrian M Caulfield - One of the best experts on this subject based on the ideXlab platform.
-
onyx a protoype phase change memory Storage Array
USENIX conference on Hot topics in storage and file systems, 2011Co-Authors: Ameen D Akel, Adrian M Caulfield, Todor Mollov, Rajesh Gupta, Steven SwansonAbstract:We describe a prototype high-performance solid-state drive based on first-generation phase-change memory (PCM) devices called Onyx. Onyx has a capacity of 10 GB and connects to the host system via PCIe. We describe the internal architecture of Onyx including the PCM memory modules we constructed and the FPGA-based controller that manages them. Onyx can perform a 4 KB random read in 38 µs and sustain 191K 4 KB read IO operations per second. A 4 KB write requires 179 µs. We describe our experience tuning the Onyx system to reduce the cost of wear-leveling and increase performance. We find that Onyx out-performs a state-of-the-art flash-based SSD for small writes (< 2 KB) by between 72 and 120% and for reads of all sizes. In addition, Onyx incurs 20-51% less CPU overhead per IOP for small requests. Combined, our results demonstrate that even first-generation PCM SSDs can outperform flash-based Arrays for the irregular (and frequently read-dominated) access patterns that define many of today's "killer" Storage applications. Next generation PCM devices will widen the performance gap further and set the stage for PCM becoming a serious flash competitor in many applications.
-
moneta a high performance Storage Array architecture for next generation non volatile memories
International Symposium on Microarchitecture, 2010Co-Authors: Adrian M Caulfield, Rajesh Gupta, Joel Coburn, Todor I Mollow, Steven SwansonAbstract:Emerging non-volatile memory technologies such as phase change memory (PCM) promise to increase Storage system performance by a wide margin relative to both conventional disks and flash-based SSDs. Realizing this potential will require significant changes to the way systems interact with Storage devices as well as a rethinking of the Storage devices themselves. This paper describes the architecture of a prototype PCIe-attached Storage Array built from emulated PCM Storage called Moneta. Moneta provides a carefully designed hardware/software interface that makes issuing and completing accesses atomic. The atomic management interface, combined with hardware scheduling optimizations, and an optimized Storage stack increases performance for small, random accesses by 18x and reduces software overheads by 60%. Moneta Array sustain 2.8~GB/s for sequential transfers and 541K random 4~KB~IO operations per second (8x higher than a state-of-the-art flash-based SSD). Moneta can perform a 512-byte write in 9~us (5.6x faster than the SSD). Moneta provides a harmonic mean speedup of 2.1x and a maximum speed up of 9x across a range of file system, paging, and database workloads. We also explore trade-offs in Moneta's architecture between performance, power, memory organization, and memory latency.
Rajesh Gupta - One of the best experts on this subject based on the ideXlab platform.
-
onyx a protoype phase change memory Storage Array
USENIX conference on Hot topics in storage and file systems, 2011Co-Authors: Ameen D Akel, Adrian M Caulfield, Todor Mollov, Rajesh Gupta, Steven SwansonAbstract:We describe a prototype high-performance solid-state drive based on first-generation phase-change memory (PCM) devices called Onyx. Onyx has a capacity of 10 GB and connects to the host system via PCIe. We describe the internal architecture of Onyx including the PCM memory modules we constructed and the FPGA-based controller that manages them. Onyx can perform a 4 KB random read in 38 µs and sustain 191K 4 KB read IO operations per second. A 4 KB write requires 179 µs. We describe our experience tuning the Onyx system to reduce the cost of wear-leveling and increase performance. We find that Onyx out-performs a state-of-the-art flash-based SSD for small writes (< 2 KB) by between 72 and 120% and for reads of all sizes. In addition, Onyx incurs 20-51% less CPU overhead per IOP for small requests. Combined, our results demonstrate that even first-generation PCM SSDs can outperform flash-based Arrays for the irregular (and frequently read-dominated) access patterns that define many of today's "killer" Storage applications. Next generation PCM devices will widen the performance gap further and set the stage for PCM becoming a serious flash competitor in many applications.
-
moneta a high performance Storage Array architecture for next generation non volatile memories
International Symposium on Microarchitecture, 2010Co-Authors: Adrian M Caulfield, Rajesh Gupta, Joel Coburn, Todor I Mollow, Steven SwansonAbstract:Emerging non-volatile memory technologies such as phase change memory (PCM) promise to increase Storage system performance by a wide margin relative to both conventional disks and flash-based SSDs. Realizing this potential will require significant changes to the way systems interact with Storage devices as well as a rethinking of the Storage devices themselves. This paper describes the architecture of a prototype PCIe-attached Storage Array built from emulated PCM Storage called Moneta. Moneta provides a carefully designed hardware/software interface that makes issuing and completing accesses atomic. The atomic management interface, combined with hardware scheduling optimizations, and an optimized Storage stack increases performance for small, random accesses by 18x and reduces software overheads by 60%. Moneta Array sustain 2.8~GB/s for sequential transfers and 541K random 4~KB~IO operations per second (8x higher than a state-of-the-art flash-based SSD). Moneta can perform a 512-byte write in 9~us (5.6x faster than the SSD). Moneta provides a harmonic mean speedup of 2.1x and a maximum speed up of 9x across a range of file system, paging, and database workloads. We also explore trade-offs in Moneta's architecture between performance, power, memory organization, and memory latency.
Julien Langou - One of the best experts on this subject based on the ideXlab platform.
-
rectangular full packed format for cholesky s algorithm factorization solution and inversion
ACM Transactions on Mathematical Software, 2010Co-Authors: Fred G Gustavson, Jerzy Waśniewski, Jack Dongarra, Julien LangouAbstract:We describe a new data format for storing triangular, symmetric, and Hermitian matrices called Rectangular Full Packed Format (RFPF). The standard two-dimensional Arrays of Fortran and C (also known as full format) that are used to represent triangular and symmetric matrices waste nearly half of the Storage space but provide high performance via the use of Level 3 BLAS. Standard packed format Arrays fully utilize Storage (Array space) but provide low performance as there is no Level 3 packed BLAS. We combine the good features of packed and full Storage using RFPF to obtain high performance via using Level 3 BLAS as RFPF is a standard full-format representation. Also, RFPF requires exactly the same minimal Storage as packed the format. Each LAPACK full and/or packed triangular, symmetric, and Hermitian routine becomes a single new RFPF routine based on eight possible data layouts of RFPF. This new RFPF routine usually consists of two calls to the corresponding LAPACK full-format routine and two calls to Level 3 BLAS routines. This means no new software is required. As examples, we present LAPACK routines for Cholesky factorization, Cholesky solution, and Cholesky inverse computation in RFPF to illustrate this new work and to describe its performance on several commonly used computer platforms. Performance of LAPACK full routines using RFPF versus LAPACK full routines using the standard format for both serial and SMP parallel processing is about the same while using half the Storage. Performance gains are roughly one to a factor of 43 for serial and one to a factor of 97 for SMP parallel times faster using vendor LAPACK full routines with RFPF than with using vendor and/or reference packed routines.
-
rectangular full packed format for cholesky s algorithm factorization solution and inversion
arXiv: Mathematical Software, 2009Co-Authors: Fred G Gustavson, Jack Dongarra, Jerzy Wasniewski, Julien LangouAbstract:We describe a new data format for storing triangular, symmetric, and Hermitian matrices called RFPF (Rectangular Full Packed Format). The standard two dimensional Arrays of Fortran and C (also known as full format) that are used to represent triangular and symmetric matrices waste nearly half of the Storage space but provide high performance via the use of Level 3 BLAS. Standard packed format Arrays fully utilize Storage (Array space) but provide low performance as there is no Level 3 packed BLAS. We combine the good features of packed and full Storage using RFPF to obtain high performance via using Level 3 BLAS as RFPF is a standard full format representation. Also, RFPF requires exactly the same minimal Storage as packed format. Each LAPACK full and/or packed triangular, symmetric, and Hermitian routine becomes a single new RFPF routine based on eight possible data layouts of RFPF. This new RFPF routine usually consists of two calls to the corresponding LAPACK full format routine and two calls to Level 3 BLAS routines. This means {\it no} new software is required. As examples, we present LAPACK routines for Cholesky factorization, Cholesky solution and Cholesky inverse computation in RFPF to illustrate this new work and to describe its performance on several commonly used computer platforms. Performance of LAPACK full routines using RFPF versus LAPACK full routines using standard format for both serial and SMP parallel processing is about the same while using half the Storage. Performance gains are roughly one to a factor of 43 for serial and one to a factor of 97 for SMP parallel times faster using vendor LAPACK full routines with RFPF than with using vendor and/or reference packed routines.
George T. Galyon - One of the best experts on this subject based on the ideXlab platform.
-
Impact of the ROHS directive on high-performance electronic systems
Journal of Materials Science: Materials in Electronics, 2007Co-Authors: Karl J. Puttlitz, George T. GalyonAbstract:The European Union enacted legislation, the ROHS Directive, that bans the use of lead (Pb) and several other substances in electronic products commencing July 1, 2006. The legislation recognized that in some situations no viable alternative Pb-free substitute materials are known at this time, and so provided exemptions for those cases. It was also recognized that certain electronic products, specifically servers, Storage and Storage Array systems, network infrastructure equipment and network management for telecommunication equipment referred to as high-performance electronic products, perform tasks so important to modern society that their operational integrity had to be maintained. The introduction of new and unproven materials posed a significant potential reliability risk. Accordingly, the European Commission (EC) granted an exemption permitting the continued use of Pb in solders, independent of concentration, for high-performance (H-P) equipment applications. This exemption was primarily aimed at assuring that the reliability of solder joints, particularly flip-chip solder joints is preserved. Flip-chip solder joints experience the most severe operating conditions in comparison to other applications that utilize Pb in electronic equipment. This paper briefly describes the solder-exempted H-P electronic products, their capabilities, and some typical tasks they perform. Also discussed are the major attributes that differentiate H-P electronic equipment from consumer electronics, particularly in relation to their operational and reliability requirements. Interestingly, other than the special solder exemption accorded to H-P electronic equipment, these products must meet all the other requirements for ROHS compliancy. The EC was aware that issues would surface after the legislation was enacted, so it created the Technical Advisory Committee (TAC) to review industry-generated requests for exemptions. The paper discusses three exemption requests granted by the EC that are particularly relevant to H-P electronic products. The exemptions allow the continued use of lead-bearing solder materials.
-
impact of the rohs directive on high performance electronic systems part i need for lead utilization in exempt systems
Journal of Materials Science: Materials in Electronics, 2006Co-Authors: Karl J. Puttlitz, George T. GalyonAbstract:The European Union enacted legislation, the ROHS Directive, that bans the use of lead (Pb) and several other substances in electronic products commencing July 1, 2006. The legislation recognized that in some situations no viable alternative Pb-free substitute materials are known at this time, and so provided exemptions for those cases. It was also recognized that certain electronic products, specifically servers, Storage and Storage Array systems, network infrastructure equipment and network management for telecommunication equipment referred to as high-performance electronic products, perform tasks so important to modern society that their operational integrity had to be maintained. The introduction of new and unproven materials posed a significant potential reliability risk. Accordingly, the European Commission (EC) granted an exemption permitting the continued use of Pb in solders, independent of concentration, for high-performance (H-P) equipment applications. This exemption was primarily aimed at assuring that the reliability of solder joints, particularly flip-chip solder joints is preserved. Flip-chip solder joints experience the most severe operating conditions in comparison to other applications that utilize Pb in electronic equipment. This paper briefly describes the solder-exempted H-P electronic products, their capabilities, and some typical tasks they perform. Also discussed are the major attributes that differentiate H-P electronic equipment from consumer electronics, particularly in relation to their operational and reliability requirements. Interestingly, other than the special solder exemption accorded to H-P electronic equipment, these products must meet all the other requirements for ROHS compliancy. The EC was aware that issues would surface after the legislation was enacted, so it created the Technical Advisory Committee (TAC) to review industry-generated requests for exemptions. The paper discusses three exemption requests granted by the EC that are particularly relevant to H-P electronic products. The exemptions allow the continued use of lead-bearing solder materials.