The Experts below are selected from a list of 867 Experts worldwide ranked by ideXlab platform
Hongyang Sun - One of the best experts on this subject based on the ideXlab platform.
-
Assessing the Impact of Partial Verifications Against Silent Data Corruptions
2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic Checkpointing approaches devised for fail-stop errors. Instead, Checkpointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we investigate the use of partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light-cost but less precise verification type in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Simulations based on a wide range of parameters confirm the benefits of partial verifications in certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
ICPP - Assessing the Impact of Partial Verifications against Silent Data Corruptions
2015 44th International Conference on Parallel Processing, 2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic check pointing approaches devised for fail-stop errors. Instead, check pointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
Assessing the impact of partial verifications against silent data corruptions
2015Co-Authors: Aurelien Cavelan, Yves Robert, Saurabh Raina, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic check pointing approaches devised for fail-stop errors. Instead, check pointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
Partial Verifications Against Silent Data Corruptions
2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Abstract: Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic Checkpointing approaches de-vised for fail-stop errors. Instead, Checkpointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pat-tern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios
Javanainen Matti - One of the best experts on this subject based on the ideXlab platform.
-
Simulations of POPC/POPS membranes with KCl
2018Co-Authors: Javanainen MattiAbstract:Simulations of POPC/POPS (5:1, 144 lipids in total) membranes with 500, 1000, 2000, 3000, and 4000 mM of KCl. Simulations of the same membrane in the presence of only counterions as well as in the presence of different amounts of CaCl_2 are available in DOI:10.5281/zenodo.897467 The systems contain 40 waters per lipid, and additional potassium counter ions for the anionic PS. Final concentrations differ from the initial as ions adsorb to the membrane. Concentration calculated as #ion/#water = [ion]/[water], where [water]=55.5 M. The lipid model by Maciejewski and Rog [1,2,3] was used. Topologies (.itp) were obtained from [3]. TIP3P water was used, and the ions were modelled using default OPLS ion parameters. All simulations are 200 ns long. Simulation parameters common for all systems are given in the md.mdp File. Simulations were performed with Gromacs 2016.3 [4]. For each KCl concentration, the trajectory (.xtc), energy File (.edr), Checkpoint File (.cpt), run input File (.tpr), index File (.ndx), topology File (.top), and the final structure (gro) are provided. [1] Maciejewski et al., J. Phys Chem. B 118, 2014, pp. 4571–4581, DOI: 10.1021/jp5016627 [2] Kulig et al., Data in Brief 5, 2015, pp. 333–336, DOI: 10.1016/j.dib.2015.09.013 [3] Rog et al., Data in Brief 7, 2016, pp. 1171–1174, DOI: 10.1016/j.dib.2016.03.067 [4] Abraham et al., SoftwareX 1–2, 2015, pp. 19–25, DOI: 10.1016/j.softx.2015.06.001
-
Simulations of POPC/POPS membranes with CaCl_2.
2018Co-Authors: Javanainen MattiAbstract:Simulations of POPC/POPS (5:1, 144 lipids in total) membranes with 0, 100, 300, 1000, and 3000 mM of CaCl_2. Additionally, the systems contain 40 waters per lipid, and potassium counter ions for the anionic PS. Final concentrations differ from the initial as ions adsorb to the membrane. Concentration calculated as #ion/#water = [ion]/[water], where [water]=55M. The lipid model by Maciejewski and Rog [1,2,3] was used. Topologies (.itp) were obtained from [3]. TIP3P water was used, and the ions were modelled using default OPLS ion parameters. All simulations are 200 ns long. Simulation parameters common for all systems are given in the .mdp File. Simulations were performed with Gromacs 5.1.4 [4] For each CaCl_2 concentration, the trajectory (.xtc), energy File (.edr), Checkpoint File (.cpt), run input File (.tpr), index File (.ndx), topology File (.top), and the final structure (gro) are provided. [1] Maciejewski et al., J. Phys Chem. B 118, 2014, pp. 4571–4581, DOI: 10.1021/jp5016627 [2] Kulig et al., Data in Brief 5, 2015, pp. 333–336, DOI: 10.1016/j.dib.2015.09.013 [3] Rog et al., Data in Brief 7, 2016, pp. 1171–1174, DOI: 10.1016/j.dib.2016.03.067 [4] Abraham et al., SoftwareX 1–2, 2015, pp. 19–25, DOI: 10.1016/j.softx.2015.06.001
-
Simulations of POPC/POPS membranes with KCl
2018Co-Authors: Javanainen MattiAbstract:Simulations of POPC/POPS (5:1, 144 lipids in total) membranes with 500, 1000, 2000, and 3000 mM of KCl. This is an update to a previous upload in which the PS tails were reversed, so it was actually 'OPPS'. Moreover, at 4000 mM of KCl the ions formed a crystal, so these data are omitted. The starting configurations were generated by CHARMM-GUI. The Files of the old submission are called "PCPS-KClx" with 'x' indicating the concentration of KCl, whereas the new and corrected Files are called "PCPC_KCL_x. Simulations of the same membrane in the presence of only counterions as well as in the presence of different amounts of CaCl_2 are available in DOI:10.5281/zenodo.897467 The systems contain 40 waters per lipid, and additional potassium counter ions for the anionic PS. Final concentrations differ from the initial as ions adsorb to the membrane. Concentration calculated as #ion/#water = [ion]/[water], where [water]=55.5 M. The lipid model by Maciejewski and Rog [1,2,3] was used. Topologies (.itp) were obtained from [3] and the OPPS chains were reversed (the fixed itp is included in the upload) to create POPS. TIP3P water was used, and the ions were modelled using default OPLS ion parameters. All simulations are 300 ns long and the uploaded trajectories contain the last 200 ns. Simulation parameters common for all systems are given in the md.mdp File. Simulations were performed with Gromacs 2016.3 [4]. For each KCl concentration, the trajectory (.xtc), energy File (.edr), Checkpoint File (.cpt), run input File (.tpr), index File (.ndx), topology File (.top), and the final structure (gro) are provided. [1] Maciejewski et al., J. Phys Chem. B 118, 2014, pp. 4571–4581, DOI: 10.1021/jp5016627 [2] Kulig et al., Data in Brief 5, 2015, pp. 333–336, DOI: 10.1016/j.dib.2015.09.013 [3] Rog et al., Data in Brief 7, 2016, pp. 1171–1174, DOI: 10.1016/j.dib.2016.03.067 [4] Abraham et al., SoftwareX 1–2, 2015, pp. 19–25, DOI: 10.1016/j.softx.2015.06.001
-
Simulation of a POPS bilayer
2017Co-Authors: Javanainen MattiAbstract:Simulation of a POPS lipid bilayer. The system contains 128 POPS lipids, 40 waters per lipid, and 128 sodium counter ions for the anionic PS. The lipid model by Maciejewski and Rog [1,2,3] was used. Topologies (.itp) were obtained from [3]. TIP3P water was used, and the sodium was modelled using default OPLS ion parameters. The simulation is 200 ns long. Simulation parameters are given in the .mdp File. Simulations were performed using Gromacs 5.1.4 [4] The trajectory (.xtc), energy File (.edr), Checkpoint File (.cpt), run input File (.tpr), index File (.ndx), topology File (.top), and the final structure (.gro) are provided. [1] Maciejewski et al., J. Phys Chem. B 118, 2014, pp. 4571–4581, DOI: 10.1021/jp5016627 [2] Kulig et al., Data in Brief 5, 2015, pp. 333–336, DOI: 10.1016/j.dib.2015.09.013 [3] Rog et al., Data in Brief 7, 2016, pp. 1171–1174, DOI: 10.1016/j.dib.2016.03.067 [4] Abraham et al., SoftwareX 1–2, 2015, pp. 19–25, DOI: 10.1016/j.softx.2015.06.00
-
Simulations of DPPC/Cholesterol bilayers with the Slipids force field, part 1/2
2017Co-Authors: Javanainen Matti, Martinez-seara Hector, Vattulainen IlpoAbstract:Simulation data related to our publication "Nanoscale Membrane Domain Formation Driven by Cholesterol" (DOI:10.1038/s41598-017-01247-9), part 1/2. Part 2/2 of the data are in the Zenodo record (DOI:10.5281/zenodo.439080). Please note that the description below covers both parts. Data for both cholesterol-free calibration simulations (letters in Table S1 and Fig. 1) and for cholesterol-containing simulations (numbers in Table S1 and Fig. 1) are included. The Files are named as "dppc-X-Y.Z", where X is the cholesterol concentration, Y the simulation temperature (not shifted, see the paper), and Z defines the File type (in GROMACS formats): xtc for trajectory, edr for energy File, tpr for the simulation input File, and .cpt for the Checkpoint File. The index File (.ndx) and the topology File (.top) are common among systems with equal cholesterol concentration. The simulation parameter File (.mdp) is common for all simulations, only the target temperature of the thermostat needs to be adjusted. Simulation lengths vary between 300 and 1400 ns, and the trajectories are written every 100 ps. Further information on the setup and composition of the simulated systems is available in the paper. The Slipids force field [1,2,3] is employed, and the topologies (.itp) are available at http://www.fos.su.se/~sasha/SLipids/ All simulations were performed with Gromacs 4.6.x [1] Derivation and Systematic Validation of a Refined All-Atom Force Field for Phosphatidylcholine Lipids. Joakim P. M. Jämbeck and Alexander P. Lyubartsev, The Journal of Physical Chemistry B 2012 116 (10), 3164-3179, DOI: 10.1021/jp212503e [2] An Extension and Further Validation of an All-Atomistic Force Field for Biological Membranes. Joakim P. M. Jämbeck and Alexander P. Lyubartsev, Journal of Chemical Theory and Computation 2012 8 (8), 2938-2948, DOI: 10.1021/ct300342n [3] Another Piece of the Membrane Puzzle: Extending Slipids Further. Joakim P. M. Jämbeck and Alexander P. Lyubartsev, Journal of Chemical Theory and Computation 2013 9 (1), 774-784, DOI: 10.1021/ct300777
Aurelien Cavelan - One of the best experts on this subject based on the ideXlab platform.
-
Assessing the Impact of Partial Verifications Against Silent Data Corruptions
2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic Checkpointing approaches devised for fail-stop errors. Instead, Checkpointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we investigate the use of partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light-cost but less precise verification type in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Simulations based on a wide range of parameters confirm the benefits of partial verifications in certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
ICPP - Assessing the Impact of Partial Verifications against Silent Data Corruptions
2015 44th International Conference on Parallel Processing, 2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic check pointing approaches devised for fail-stop errors. Instead, check pointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
Assessing the impact of partial verifications against silent data corruptions
2015Co-Authors: Aurelien Cavelan, Yves Robert, Saurabh Raina, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic check pointing approaches devised for fail-stop errors. Instead, check pointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
Partial Verifications Against Silent Data Corruptions
2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Abstract: Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic Checkpointing approaches de-vised for fail-stop errors. Instead, Checkpointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pat-tern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios
Maria Martin - One of the best experts on this subject based on the ideXlab platform.
-
improving an mpi application level migration approach through Checkpoint File splitting
Symposium on Computer Architecture and High Performance Computing, 2014Co-Authors: Monica Rodriguez, Ivan Cores, Patricia Gonzalez, Maria MartinAbstract:Traditionally used for load balancing, process migration has been gaining popularity in the fault tolerance context. Recently, Checkpoint-based migration has been proposed to implement failure avoidance in MPI applications through the proactive migration of processes when impending failures are notified. However, the main drawback of Checkpoint-based migration in these scenarios is its high I/0 cost, which may be unfeasible if the migration operation is not completed before the failure arises. To overcome this issue, this work proposes to split the Checkpoint Files of an application-level migration approach into multiple smaller Files to overlap the different phase of the migration operation: Checkpoint File writing in the terminating process, with data transferring through the network, and state File read and restart operations in the new spawned processes. The proposal has been tested using the MPI NAS Parallel Benchmarks. The experimental results show a significant reduction in the migration time.
-
SBAC-PAD - Improving an MPI Application-Level Migration Approach through Checkpoint File Splitting
2014 IEEE 26th International Symposium on Computer Architecture and High Performance Computing, 2014Co-Authors: Monica Rodriguez, Ivan Cores, Patricia Gonzalez, Maria MartinAbstract:Traditionally used for load balancing, process migration has been gaining popularity in the fault tolerance context. Recently, Checkpoint-based migration has been proposed to implement failure avoidance in MPI applications through the proactive migration of processes when impending failures are notified. However, the main drawback of Checkpoint-based migration in these scenarios is its high I/0 cost, which may be unfeasible if the migration operation is not completed before the failure arises. To overcome this issue, this work proposes to split the Checkpoint Files of an application-level migration approach into multiple smaller Files to overlap the different phase of the migration operation: Checkpoint File writing in the terminating process, with data transferring through the network, and state File read and restart operations in the new spawned processes. The proposal has been tested using the MPI NAS Parallel Benchmarks. The experimental results show a significant reduction in the migration time.
-
Extending an Application-Level Checkpointing Tool to Provide Fault Tolerance Support to OpenMP Applications
Journal of Universal Computer Science, 2014Co-Authors: Nuria Losada, Gabriel Rodriguez, Maria Martin, Patricia GonzalezAbstract:Despite the increasing popularity of shared-memory systems, there is a lack of tools for providing fault tolerance support to shared-memory applications. CPPC (ComPiler for Portable Checkpointing) is an application-level Checkpointing tool fo- cused on the insertion of fault tolerance into long-running MPI applications. This paper presents an extension to CPPC to allow the Checkpointing of OpenMP applica- tions. The proposed solution maintains the main characteristics of CPPC: portability and reduced Checkpoint File size. The performance of the proposal is evaluated using the OpenMP NAS Parallel Benchmarks showing that most of the applications present small Checkpoint overheads.
-
reducing application level Checkpoint File sizes towards scalable fault tolerance solutions
International Symposium on Parallel and Distributed Processing and Applications, 2012Co-Authors: Ivn Cores, Gabriel Rodriguez, Maria Martin, Patricia GonzlezAbstract:Systems intended for the execution of long-running parallel applications require fault tolerant capabilities, since the probability of failure increases with the execution time and the number of nodes. Checkpointing and rollback recovery is one of the most popular techniques to provide fault tolerance support. However, in order to be useful for large scale systems, current Checkpoint-recovery techniques should tackle the problem of reducing Checkpointing cost. This paper addresses this issue through the reduction of the Checkpoint File sizes. Different solutions to reduce the size of the Checkpoints generated at application level are proposed and implemented in a Checkpointing tool. Detailed experimental results on two multicore clusters show the effectiveness of the proposed methods.
-
ISPA - Reducing Application-level Checkpoint File Sizes: Towards Scalable Fault Tolerance Solutions
2012 IEEE 10th International Symposium on Parallel and Distributed Processing with Applications, 2012Co-Authors: Ivn Cores, Gabriel Rodriguez, Maria Martin, Patricia Gonz'lezAbstract:Systems intended for the execution of long-running parallel applications require fault tolerant capabilities, since the probability of failure increases with the execution time and the number of nodes. Checkpointing and rollback recovery is one of the most popular techniques to provide fault tolerance support. However, in order to be useful for large scale systems, current Checkpoint-recovery techniques should tackle the problem of reducing Checkpointing cost. This paper addresses this issue through the reduction of the Checkpoint File sizes. Different solutions to reduce the size of the Checkpoints generated at application level are proposed and implemented in a Checkpointing tool. Detailed experimental results on two multicore clusters show the effectiveness of the proposed methods.
Yves Robert - One of the best experts on this subject based on the ideXlab platform.
-
Assessing the Impact of Partial Verifications Against Silent Data Corruptions
2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic Checkpointing approaches devised for fail-stop errors. Instead, Checkpointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we investigate the use of partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light-cost but less precise verification type in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Simulations based on a wide range of parameters confirm the benefits of partial verifications in certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
ICPP - Assessing the Impact of Partial Verifications against Silent Data Corruptions
2015 44th International Conference on Parallel Processing, 2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic check pointing approaches devised for fail-stop errors. Instead, check pointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
Assessing the impact of partial verifications against silent data corruptions
2015Co-Authors: Aurelien Cavelan, Yves Robert, Saurabh Raina, Hongyang SunAbstract:Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic check pointing approaches devised for fail-stop errors. Instead, check pointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pattern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios, when compared to the baseline algorithm that uses only guaranteed verifications.
-
Partial Verifications Against Silent Data Corruptions
2015Co-Authors: Aurelien Cavelan, Saurabh Kumar Raina, Yves Robert, Hongyang SunAbstract:Abstract: Silent errors, or silent data corruptions, constitute a major threat on very large scale platforms. When a silent error strikes, it is not detected immediately but only after some delay, which prevents the use of pure periodic Checkpointing approaches de-vised for fail-stop errors. Instead, Checkpointing must be coupled with some verification mechanism to guarantee that corrupted data will never be written into the Checkpoint File. Such a guaranteed verification mechanism typically incurs a high cost. In this paper, we assess the impact of using partial verification mechanisms in addition to a guaranteed verification. The main objective is to investigate to which extent it is worthwhile to use some light cost but less accurate verifications in the middle of a periodic computing pat-tern, which ends with a guaranteed verification right before each Checkpoint. Introducing partial verifications dramatically complicates the analysis, but we are able to analytically determine the optimal computing pattern (up to the first-order approximation), including the optimal length of the pattern, the optimal number of partial verifications, as well as their optimal positions inside the pattern. Performance evaluations based on a wide range of parameters confirm the benefit of using partial verifications under certain scenarios