The Experts below are selected from a list of 62928 Experts worldwide ranked by ideXlab platform
Milind Kulkarni - One of the best experts on this subject based on the ideXlab platform.
-
automatically enhancing locality for tree traversals with traversal splicing
Conference on Object-Oriented Programming Systems Languages and Applications, 2012Co-Authors: Milind KulkarniAbstract:Generally applicable techniques for improving temporal locality in irregular programs, which operate over pointer-based data structures such as trees and graphs, are scarce. Focusing on a subset of irregular programs, namely, tree traversal algorithms like Barnes-Hut and nearest neighbor, previous work has proposed point blocking, a technique analogous to loop tiling in regular programs, to improve locality. However point blocking is highly dependent on point sorting, a technique to reorder points so that consecutive points will have similar traversals. Performing this a priori sort requires an understanding of the semantics of the algorithm and hence highly application specific techniques. In this work, we propose traversal splicing, a new, general, automatic locality optimization for irregular tree traversal codes, that is less sensitive to point order, and hence can deliver substantially better performance, even in the absence of semantic information. For six benchmark algorithms, we show that traversal splicing can deliver single-thread speedups of up to 9.147 (geometric mean: 3.095) over Baseline Implementations, and up to 4.752 (geometric mean: 2.079) over point-blocked Implementations. Further, we show that in many cases, automatically applying traversal splicing to a Baseline Implementation yields performance that is better than carefully hand-optimized Implementations.
-
brief announcement locality enhancing loop transformations for tree traversal algorithms
ACM Symposium on Parallel Algorithms and Architectures, 2011Co-Authors: Milind KulkarniAbstract:In this paper, we discuss transformations that can be applied to irregular programs that perform tree traversals, which can be seen as analogs of the popular regular transformations of loop tiling. We demonstrate the utility of these transformations on two tree traversal algorithms, the Barnes-Hut algorithm and raytracing, achieving speedups of up to 237% over the Baseline Implementation.
-
locality enhancing loop transformations for parallel tree traversal algorithms
2011Co-Authors: Milind KulkarniAbstract:Exploiting locality is critical to achieving good performance. For regular programs, which operate on dense arrays and matrices, techniques such as loop interchange and tiling have long been known to improve locality and deliver improved performance. However, there has been relatively little work investigating similar locality-improving transformations for irregular programs that operate on trees or graphs. Often, it is not even clear that such transformations are possible. In this paper, we discuss two transformations that can be applied to irregular programs that perform graph traversals. We show that these transformations can be seen as analogs of the popular regular transformations of loop interchange and tiling. We demonstrate the utility of these transformations on two tree traversal algorithms, the Barnes-Hut algorithm and raytracing, achieving speedups of up to 251% over the Baseline Implementation.
Luca Benini - One of the best experts on this subject based on the ideXlab platform.
-
energy efficient hardware accelerated synchronization for shared l1 memory multiprocessor clusters
IEEE Transactions on Parallel and Distributed Systems, 2021Co-Authors: Florian Glaser, Giuseppe Tagliavini, Davide Rossi, Germain Haugou, Qiuting Huang, Luca BeniniAbstract:The steeply growing performance demands for highly power- and energy-constrained processing systems such as end-nodes of the Internet-of-Things (IoT) have led to parallel near-threshold computing (NTC), joining the energy-efficiency benefits of low-voltage operation with the performance typical of parallel systems. Shared-L1-memory multiprocessor clusters are a promising architecture, delivering performance in the order of GOPS and over 100 GOPS/W of energy-efficiency. However, this level of computational efficiency can only be reached by maximizing the effective utilization of the processing elements (PEs) available in the clusters. Along with this effort, the optimization of PE-to-PE synchronization and communication is a critical factor for performance. In this article, we describe a light-weight hardware-accelerated synchronization and communication unit (SCU) for tightly-coupled clusters of processors. We detail the architecture, which enables fine-grain per-PE power management, and its integration into an eight-core cluster of RISC-V processors. To validate the effectiveness of the proposed solution, we implemented the eight-core cluster in advanced 22 nm FDX technology and evaluated performance and energy-efficiency with tunable microbenchmarks and a set of real-life applications and kernels. The proposed solution allows synchronization-free regions as small as 42 cycles, over 41× smaller than the Baseline Implementation based on fast test-and-set access to L1 memory when constraining the microbenchmarks to 10 percent synchronization overhead. When evaluated on the real-life DSP-applications, the proposed SCU improves performance by up to 92 and 23 percent on average and energy efficiency by up to 98 and 39 percent on average.
-
pulp nn a computing library for quantized neural network inference at the edge on risc v based parallel ultra low power clusters
International Conference on Electronics Circuits and Systems, 2019Co-Authors: Angelo Garofalo, Davide Rossi, Manuele Rusci, Francesco Conti, Luca BeniniAbstract:We present PULP-NN, a multicore computing library for a parallel ultra-low-power cluster of RISC-V based processors. The library consists of a set of kernels for Quantized Neural Network (QNN) inference on edge devices, targeting byte and sub-byte data types, down to INT-1. Our software solution exploits the digital signal processing (DSP) extensions available in the PULP RISC-V processors and the cluster's parallelism, improving performance by up to 63× with respect to a Baseline Implementation on a single RISC-V core implementing the RV32IMC ISA. Using the PULP-NN routines, the inference of a CIFAR-10 QNN model runs in 30× and 19.6× less clock cycles than the current state-of-the-art ARM CMSIS-NN library, running on an STM32L4 and an STM32H7 MCUs, respectively. By running the library kernels on the GAP-8 processor at the maximum efficiency operating point, the energy efficiency on GAP-8 is 14.1× higher than STM32L4 and 39.5× than STM32H7.
Brendan Mcgrath - One of the best experts on this subject based on the ideXlab platform.
-
improving tracheostomy care in the united kingdom results of a guided quality improvement programme in 20 diverse hospitals
BJA: British Journal of Anaesthesia, 2020Co-Authors: Brendan Mcgrath, Sarah E Wallace, James P Lynch, Barbara Bonvento, Barry Coe, Anna Owen, Mike Firn, Michael Brenner, Elizabeth A EdwardsAbstract:Abstract Background Inconsistent and poorly coordinated systems of tracheostomy care commonly result in frustrations, delays, and harm. Quality improvement strategies described by exemplar hospitals of the Global Tracheostomy Collaborative have potential to mitigate such problems. This 3 yr guided Implementation programme investigated interventions designed to improve the quality and safety of tracheostomy care. Methods The programme management team guided the Implementation of 18 interventions over three phases (Baseline/Implementation/evaluation). Mixed-methods interviews, focus groups, and Hospital Anxiety and Depression Scale questionnaires defined outcome measures, with patient-level databases tracking and benchmarking process metrics. Appreciative inquiry, interviews, and Normalisation Measure Development questionnaires explored change barriers and enablers. Results All sites implemented at least 16/18 interventions, with the magnitude of some improvements linked to staff engagement (1536 questionnaires from 1019 staff), and 2405 admissions (1868 ICU/high-dependency unit; 7.3% children) were prospectively captured. Median stay was 50 hospital days, 23 ICU days, and 28 tracheostomy days. Incident severity score reduced significantly (n=606; P Conclusions This guided improvement programme for tracheostomy patients significantly improved the quality and safety of care, contributing rich qualitative improvement data. Patient-centred outcomes were improved along with significant efficiency and cost savings across diverse UK hospitals. Clinical trial registration IRAS-ID-206955; REC-Ref-16/LO/1196; NIHR Portfolio CPMS ID 31544.
Florian Glaser - One of the best experts on this subject based on the ideXlab platform.
-
energy efficient hardware accelerated synchronization for shared l1 memory multiprocessor clusters
IEEE Transactions on Parallel and Distributed Systems, 2021Co-Authors: Florian Glaser, Giuseppe Tagliavini, Davide Rossi, Germain Haugou, Qiuting Huang, Luca BeniniAbstract:The steeply growing performance demands for highly power- and energy-constrained processing systems such as end-nodes of the Internet-of-Things (IoT) have led to parallel near-threshold computing (NTC), joining the energy-efficiency benefits of low-voltage operation with the performance typical of parallel systems. Shared-L1-memory multiprocessor clusters are a promising architecture, delivering performance in the order of GOPS and over 100 GOPS/W of energy-efficiency. However, this level of computational efficiency can only be reached by maximizing the effective utilization of the processing elements (PEs) available in the clusters. Along with this effort, the optimization of PE-to-PE synchronization and communication is a critical factor for performance. In this article, we describe a light-weight hardware-accelerated synchronization and communication unit (SCU) for tightly-coupled clusters of processors. We detail the architecture, which enables fine-grain per-PE power management, and its integration into an eight-core cluster of RISC-V processors. To validate the effectiveness of the proposed solution, we implemented the eight-core cluster in advanced 22 nm FDX technology and evaluated performance and energy-efficiency with tunable microbenchmarks and a set of real-life applications and kernels. The proposed solution allows synchronization-free regions as small as 42 cycles, over 41× smaller than the Baseline Implementation based on fast test-and-set access to L1 memory when constraining the microbenchmarks to 10 percent synchronization overhead. When evaluated on the real-life DSP-applications, the proposed SCU improves performance by up to 92 and 23 percent on average and energy efficiency by up to 98 and 39 percent on average.
Elizabeth A Edwards - One of the best experts on this subject based on the ideXlab platform.
-
improving tracheostomy care in the united kingdom results of a guided quality improvement programme in 20 diverse hospitals
BJA: British Journal of Anaesthesia, 2020Co-Authors: Brendan Mcgrath, Sarah E Wallace, James P Lynch, Barbara Bonvento, Barry Coe, Anna Owen, Mike Firn, Michael Brenner, Elizabeth A EdwardsAbstract:Abstract Background Inconsistent and poorly coordinated systems of tracheostomy care commonly result in frustrations, delays, and harm. Quality improvement strategies described by exemplar hospitals of the Global Tracheostomy Collaborative have potential to mitigate such problems. This 3 yr guided Implementation programme investigated interventions designed to improve the quality and safety of tracheostomy care. Methods The programme management team guided the Implementation of 18 interventions over three phases (Baseline/Implementation/evaluation). Mixed-methods interviews, focus groups, and Hospital Anxiety and Depression Scale questionnaires defined outcome measures, with patient-level databases tracking and benchmarking process metrics. Appreciative inquiry, interviews, and Normalisation Measure Development questionnaires explored change barriers and enablers. Results All sites implemented at least 16/18 interventions, with the magnitude of some improvements linked to staff engagement (1536 questionnaires from 1019 staff), and 2405 admissions (1868 ICU/high-dependency unit; 7.3% children) were prospectively captured. Median stay was 50 hospital days, 23 ICU days, and 28 tracheostomy days. Incident severity score reduced significantly (n=606; P Conclusions This guided improvement programme for tracheostomy patients significantly improved the quality and safety of care, contributing rich qualitative improvement data. Patient-centred outcomes were improved along with significant efficiency and cost savings across diverse UK hospitals. Clinical trial registration IRAS-ID-206955; REC-Ref-16/LO/1196; NIHR Portfolio CPMS ID 31544.