The Experts below are selected from a list of 91077 Experts worldwide ranked by ideXlab platform
Dong H. Ahn - One of the best experts on this subject based on the ideXlab platform.
-
HPDC - Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters
Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing, 2016Co-Authors: Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas R. W. Scogland, Marc D. Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela TauferAbstract:The economics of flash vs. disk storage is driving HPC centers to incorporate faster solid-state burst buffers into the storage hierarchy in exchange for smaller parallel file system (PFS) bandwidth. In systems with an underprovisioned PFS, avoiding I/O contention at the PFS level will become crucial to achieving high computational efficiency. In this paper, we propose novel batch job scheduling techniques that reduce such contention by integrating I/O awareness into scheduling policies such as EASY backfilling. We model the available bandwidth of links between each level of the storage hierarchy (i.e., burst buffers, I/O network, and PFS), and our I/O-aware schedulers use this model to avoid contention at any level in the hierarchy. We integrate our approach into Flux, a next-Generation Resource and job management framework, and evaluate the effectiveness and computational costs of our I/O-aware scheduling. Our results show that by reducing I/O contention for underprovisioned PFSes, our solution reduces job performance variability by up to 33% and decreases I/O-related utilization losses by up to 21%, which ultimately increases the amount of science performed by scientific workloads.
-
flux a next Generation Resource management framework for large hpc centers
International Conference on Parallel Processing, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
-
ICPP Workshops - Flux: A Next-Generation Resource Management Framework for Large HPC Centers
2014 43rd International Conference on Parallel Processing Workshops, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
Mohammad Shahidehpour - One of the best experts on this subject based on the ideXlab platform.
-
Stochastic Midterm Coordination of Hydro and Natural Gas Flexibilities for Wind Energy Integration
IEEE Transactions on Sustainable Energy, 2014Co-Authors: Saeed Kamalinia, Lei Wu, Mohammad ShahidehpourAbstract:This paper presents a stochastic security-constrained unit commitment (SCUC) model for the optimal midterm and flexible allocations of hydro and natural gas systems when accommodating a large integration of wind energy. Random errors in forecasting the hourly wind, natural water inflow, and electric load as well as random outages of power system components are modeled in scenarios via the Monte Carlo (MC) simulation. The proposed optimization model is formulated as a two-stage stochastic problem, where short-term and midterm Generation Resource optimizations are investigated in the first and the second stages, respectively. The reliability criteria including the loss of load expectation (LOLE) and load curtailment limits are incorporated into the midterm stochastic SCUC problem, where electric power and natural gas network constraints are checked. Numerical experiments signify the effectiveness of the proposed method for the midterm optimal scheduling of water and natural gas flexibilities when integrating wind energy Resources.
-
security constrained Resource planning in electricity markets
IEEE Transactions on Power Systems, 2007Co-Authors: Jae Hyung Roh, Mohammad ShahidehpourAbstract:We propose a security-based competitive Generation Resource planning model in electricity markets. The objective of the model is to introduce the impact of transmission security in a multi-GENCO Generation Resource planning. The proposed approach is based on effective decomposition and coordination strategies. Lagrangian relaxation and Benders decomposition are applied to decompose the original planning problem into tractable optimization subproblems. Locational price signal and security signal are defined for the simulation of competition among GENCOs and the coordination of security between GENCOs and the regulatory body (ISO). The numerical examples exhibit the effectiveness of the proposed Generation planning model in electricity markets
Martin Schulz - One of the best experts on this subject based on the ideXlab platform.
-
flux a next Generation Resource management framework for large hpc centers
International Conference on Parallel Processing, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
-
ICPP Workshops - Flux: A Next-Generation Resource Management Framework for Large HPC Centers
2014 43rd International Conference on Parallel Processing Workshops, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
Becky Springmeyer - One of the best experts on this subject based on the ideXlab platform.
-
HPDC - Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters
Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing, 2016Co-Authors: Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas R. W. Scogland, Marc D. Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela TauferAbstract:The economics of flash vs. disk storage is driving HPC centers to incorporate faster solid-state burst buffers into the storage hierarchy in exchange for smaller parallel file system (PFS) bandwidth. In systems with an underprovisioned PFS, avoiding I/O contention at the PFS level will become crucial to achieving high computational efficiency. In this paper, we propose novel batch job scheduling techniques that reduce such contention by integrating I/O awareness into scheduling policies such as EASY backfilling. We model the available bandwidth of links between each level of the storage hierarchy (i.e., burst buffers, I/O network, and PFS), and our I/O-aware schedulers use this model to avoid contention at any level in the hierarchy. We integrate our approach into Flux, a next-Generation Resource and job management framework, and evaluate the effectiveness and computational costs of our I/O-aware scheduling. Our results show that by reducing I/O contention for underprovisioned PFSes, our solution reduces job performance variability by up to 33% and decreases I/O-related utilization losses by up to 21%, which ultimately increases the amount of science performed by scientific workloads.
-
flux a next Generation Resource management framework for large hpc centers
International Conference on Parallel Processing, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
-
ICPP Workshops - Flux: A Next-Generation Resource Management Framework for Large HPC Centers
2014 43rd International Conference on Parallel Processing Workshops, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
Jim Garlick - One of the best experts on this subject based on the ideXlab platform.
-
HPDC - Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters
Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing, 2016Co-Authors: Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas R. W. Scogland, Marc D. Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela TauferAbstract:The economics of flash vs. disk storage is driving HPC centers to incorporate faster solid-state burst buffers into the storage hierarchy in exchange for smaller parallel file system (PFS) bandwidth. In systems with an underprovisioned PFS, avoiding I/O contention at the PFS level will become crucial to achieving high computational efficiency. In this paper, we propose novel batch job scheduling techniques that reduce such contention by integrating I/O awareness into scheduling policies such as EASY backfilling. We model the available bandwidth of links between each level of the storage hierarchy (i.e., burst buffers, I/O network, and PFS), and our I/O-aware schedulers use this model to avoid contention at any level in the hierarchy. We integrate our approach into Flux, a next-Generation Resource and job management framework, and evaluate the effectiveness and computational costs of our I/O-aware scheduling. Our results show that by reducing I/O contention for underprovisioned PFSes, our solution reduces job performance variability by up to 33% and decreases I/O-related utilization losses by up to 21%, which ultimately increases the amount of science performed by scientific workloads.
-
flux a next Generation Resource management framework for large hpc centers
International Conference on Parallel Processing, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.
-
ICPP Workshops - Flux: A Next-Generation Resource Management Framework for Large HPC Centers
2014 43rd International Conference on Parallel Processing Workshops, 2014Co-Authors: Dong H. Ahn, Don Lipari, Mark Grondona, Jim Garlick, Becky Springmeyer, Martin SchulzAbstract:Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, Resource and workload diversity, interplays between various Resources (e.g., between compute clusters and a global file system), and complexity of Resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and Resource management framework specifically designed to deal with the requirements of next-Generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of Resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.