The Experts below are selected from a list of 26640 Experts worldwide ranked by ideXlab platform

Mary Lou Soffa - One of the best experts on this subject based on the ideXlab platform.

  • Predicting the Memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines
    2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
    Co-Authors: Wei Wang, Jack W. Davidson, Mary Lou Soffa
    Abstract:

    Modern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available Memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA Memory Configuration and the asymmetric application Memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA Memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both Memory bandwidth usage and optimal core allocations. NuCore considers various Memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA Memory systems and predicts Memory bandwidth usages with only 10% average error.

  • HPCA - Predicting the Memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines
    2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
    Co-Authors: Wei Wang, Jack W. Davidson, Mary Lou Soffa
    Abstract:

    Modern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available Memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA Memory Configuration and the asymmetric application Memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA Memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both Memory bandwidth usage and optimal core allocations. NuCore considers various Memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA Memory systems and predicts Memory bandwidth usages with only 10% average error.

Wei Wang - One of the best experts on this subject based on the ideXlab platform.

  • Predicting the Memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines
    2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
    Co-Authors: Wei Wang, Jack W. Davidson, Mary Lou Soffa
    Abstract:

    Modern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available Memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA Memory Configuration and the asymmetric application Memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA Memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both Memory bandwidth usage and optimal core allocations. NuCore considers various Memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA Memory systems and predicts Memory bandwidth usages with only 10% average error.

  • HPCA - Predicting the Memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines
    2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
    Co-Authors: Wei Wang, Jack W. Davidson, Mary Lou Soffa
    Abstract:

    Modern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available Memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA Memory Configuration and the asymmetric application Memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA Memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both Memory bandwidth usage and optimal core allocations. NuCore considers various Memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA Memory systems and predicts Memory bandwidth usages with only 10% average error.

Wen-tsong Shiue - One of the best experts on this subject based on the ideXlab platform.

  • Multi-Module Multi-Port Memory Design for Low Power Embedded Systems
    Design Automation for Embedded Systems, 2004
    Co-Authors: Wen-tsong Shiue, Chaitali Chakrabarti
    Abstract:

    In this paper we describe a multi-module, multi-port Memory design procedure that satisfies area and/or energy constraints for embedded applications. Our procedure consists of application of loop transformations and reordering of array accesses to reduce the Memory bandwidth followed by Memory allocation and assignment procedures based on ILP models and heuristic-based algorithms. The specific problems include determination of (a) the Memory Configuration with minimum area, given the energy bound, (b) the Memory Configuration with minimum energy, given the area bound, (c) array allocation such that the energy consumption is minimum for a given Memory Configuration (number of modules, size and number of ports per module). The results obtained by the heuristics match well with those obtained by the ILP methods.

  • ASAP - Low power Memory design
    Proceedings IEEE International Conference on Application- Specific Systems Architectures and Processors, 2002
    Co-Authors: Wen-tsong Shiue
    Abstract:

    In this paper, we present a novel design procedure for multi-module, multi-port Memory design that satisfies area and/or energy/timing constraints. Our procedure consists of (i) use of storage bandwidth optimization (SBO) techniques to simplify the conflict graph and (ii) use of Memory exploration techniques to determine the best Memory Configuration (number of modules, size and number of ports per module) with the minimum area if the energy and timing are bounded or with the minimum energy/timing if the area is bounded. Here the simplest conflict graph implies more possibilities for the arrays assigned to the same module without the penalty in an increase of the number of ports for each module. Our benchmark shows that the heuristic algorithm is very efficient to decide the best Memory Configuration for the system constraints (timing, area, or energy). In addition, the CACTI tool (Premkishore Shivakumar and N.P. Jouppi, 2001) is modified to estimate the timing, area, and energy for each module in different CMOS technologies (0.8 /spl mu/m, 0.35 /spl mu/m, and 0.18 /spl mu/m). Furthermore, we consider the lifetime for arrays; this results in significant reduction in timing, area, and energy for the arrays executed in different cycles sharing the same Memory module.

  • Low power multi-module, multi-port Memory design for embedded systems
    2000 IEEE Workshop on SiGNAL PROCESSING SYSTEMS. SiPS 2000. Design and Implementation (Cat. No.00TH8528), 2000
    Co-Authors: Wen-tsong Shiue, S. Tadas, Chaitali Chakrabarti
    Abstract:

    In this paper we describe a multi-module, multi-port Memory design procedure that satisfies area and/or energy constraints. Our procedure consists of use of ILP models and heuristic-based algorithms to determine (a) the Memory Configuration with minimum area, given the energy bound, (b) the Memory Configuration with minimum energy, given the area bound, (c) array allocation such that the energy consumption is minimum for a given Memory Configuration (number of modules, size and number of ports per module). The results obtained by the heuristics match very well with those obtained by the ILP methods.

  • Minimizing area/energy for low power Memory design using integer linear programming
    Proceedings of the 43rd IEEE Midwest Symposium on Circuits and Systems (Cat.No.CH37144), 2000
    Co-Authors: Wen-tsong Shiue
    Abstract:

    In this paper we describe a multi-module, multiport Memory design procedure that satisfies area and/or energy constraints, in addition to the cycle budget. We show how loop transformations can be used to derive architectures with fewer Memory modules and fewer Memory ports. We develop ILP-based models (to obtain optimal solutions) to determine the Memory Configuration for the case when (i) area has to be minimized, given the energy bound, (ii) energy has to be minimized, given the area bound.

  • DAC - Memory exploration for low power, embedded systems
    Proceedings of the 36th ACM IEEE conference on Design automation conference - DAC '99, 1999
    Co-Authors: Wen-tsong Shiue, Chaitali Chakrabarti
    Abstract:

    In embedded system design, the designer has to choose an on-chip Memory Configuration that is suitable for a specific application. To aid in this design choice, we present a Memory exploration strategy based on three performance metrics, namely, cache size, the number of processor cycles and the energy consumption. We show how the performance is affected by cache parameters such as cache size, line size, set associativity and tiling, and the off-chip data organization. We show the importance of including energy in the performance metrics, since an increase in the cache line size, cache size, tiling and set associativity reduces the number of cycles but does not necessarily reduce the energy consumption. These performance metrics help us find the minimum energy cache Configuration if time is the hard constraint, or the minimum time cache Configuration if energy is the hard constraint.

Chaitali Chakrabarti - One of the best experts on this subject based on the ideXlab platform.

  • Multi-Module Multi-Port Memory Design for Low Power Embedded Systems
    Design Automation for Embedded Systems, 2004
    Co-Authors: Wen-tsong Shiue, Chaitali Chakrabarti
    Abstract:

    In this paper we describe a multi-module, multi-port Memory design procedure that satisfies area and/or energy constraints for embedded applications. Our procedure consists of application of loop transformations and reordering of array accesses to reduce the Memory bandwidth followed by Memory allocation and assignment procedures based on ILP models and heuristic-based algorithms. The specific problems include determination of (a) the Memory Configuration with minimum area, given the energy bound, (b) the Memory Configuration with minimum energy, given the area bound, (c) array allocation such that the energy consumption is minimum for a given Memory Configuration (number of modules, size and number of ports per module). The results obtained by the heuristics match well with those obtained by the ILP methods.

  • Low power multi-module, multi-port Memory design for embedded systems
    2000 IEEE Workshop on SiGNAL PROCESSING SYSTEMS. SiPS 2000. Design and Implementation (Cat. No.00TH8528), 2000
    Co-Authors: Wen-tsong Shiue, S. Tadas, Chaitali Chakrabarti
    Abstract:

    In this paper we describe a multi-module, multi-port Memory design procedure that satisfies area and/or energy constraints. Our procedure consists of use of ILP models and heuristic-based algorithms to determine (a) the Memory Configuration with minimum area, given the energy bound, (b) the Memory Configuration with minimum energy, given the area bound, (c) array allocation such that the energy consumption is minimum for a given Memory Configuration (number of modules, size and number of ports per module). The results obtained by the heuristics match very well with those obtained by the ILP methods.

  • DAC - Memory exploration for low power, embedded systems
    Proceedings of the 36th ACM IEEE conference on Design automation conference - DAC '99, 1999
    Co-Authors: Wen-tsong Shiue, Chaitali Chakrabarti
    Abstract:

    In embedded system design, the designer has to choose an on-chip Memory Configuration that is suitable for a specific application. To aid in this design choice, we present a Memory exploration strategy based on three performance metrics, namely, cache size, the number of processor cycles and the energy consumption. We show how the performance is affected by cache parameters such as cache size, line size, set associativity and tiling, and the off-chip data organization. We show the importance of including energy in the performance metrics, since an increase in the cache line size, cache size, tiling and set associativity reduces the number of cycles but does not necessarily reduce the energy consumption. These performance metrics help us find the minimum energy cache Configuration if time is the hard constraint, or the minimum time cache Configuration if energy is the hard constraint.

  • ISCAS (1) - Memory exploration for low power embedded systems
    ISCAS'99. Proceedings of the 1999 IEEE International Symposium on Circuits and Systems VLSI (Cat. No.99CH36349), 1
    Co-Authors: Wen-tsong Shiue, Chaitali Chakrabarti
    Abstract:

    In embedded system design, the designer has to choose an on-chip Memory Configuration that is suitable for a specific application. To aid in this design choice, we present a Memory exploration strategy based on three performance metrics, namely, cache size, the number of processor cycles and the energy consumption. We show how the performance is affected by cache parameters such as cache size, line size, set associativity and tiling, and the off-chip data organization. We show the importance of including energy in the performance metrics, since an increase in the cache line size, cache size, tiling and set associativity reduces the number of cycles but does not necessarily reduce the energy consumption.

Jack W. Davidson - One of the best experts on this subject based on the ideXlab platform.

  • Predicting the Memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines
    2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
    Co-Authors: Wei Wang, Jack W. Davidson, Mary Lou Soffa
    Abstract:

    Modern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available Memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA Memory Configuration and the asymmetric application Memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA Memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both Memory bandwidth usage and optimal core allocations. NuCore considers various Memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA Memory systems and predicts Memory bandwidth usages with only 10% average error.

  • HPCA - Predicting the Memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines
    2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
    Co-Authors: Wei Wang, Jack W. Davidson, Mary Lou Soffa
    Abstract:

    Modern NUMA platforms offer large numbers of cores to boost performance through parallelism and multi-threading. However, because performance scalability is limited by available Memory bandwidth, the strategy of allocating all cores can result in degraded performance. Consequently, accurately predicting optimal (best performing) core allocations, and executing applications with these allocations are crucial for achieving the best performance. Previous research focused on the prediction of optimal numbers of cores. However, in this paper, we show that, because of the asymmetric NUMA Memory Configuration and the asymmetric application Memory behavior, optimal core allocations are not merely optimal numbers of cores. Additionally, previous studies do not adequately consider NUMA Memory resources, which further limits their ability to accurately predict optimal core allocations. In this paper, we present a model, NuCore, which predicts both Memory bandwidth usage and optimal core allocations. NuCore considers various Memory resources and NUMA asymmetry, and employs Integer Programming to achieve high accuracy and low overhead. Experimental results from real NUMA machines show that the core allocations predicted by NuCore provide 1.27x average speedup over using all cores with only 75.6% cores allocated. NuCore also provides 1.18x and 1.21x average speedups over two state-of-the-art techniques. Our results also show that NuCore faithfully models NUMA Memory systems and predicts Memory bandwidth usages with only 10% average error.