The Experts below are selected from a list of 306 Experts worldwide ranked by ideXlab platform
Darrell D.e. Long - One of the best experts on this subject based on the ideXlab platform.
-
three dimensional redundancy codes for Archival Storage
Modeling Analysis and Simulation On Computer and Telecommunication Systems, 2013Co-Authors: J F Paris, Darrell D.e. Long, Witold LitwinAbstract:Fault-tolerant disk arrays rely on replication or erasure-coding to reconstruct lost data after a disk failure. As disk capacity increases, so does the risk of encountering irrecoverable read errors that would prevent the full recovery of the lost data. We propose a three-dimensional erasure-coding technique that reduces that risk by guaranteeing full recovery in the presence of all triple and nearly all quadruple disk failures. Our solution performs better than existing solutions, such as sets of disk arrays using Reed-Solomon codes against triple failures in each individual array. Given its very high reliability, it is especially suited to the needs of very large data sets that must be preserved over long periods of time.
-
MASCOTS - Three-Dimensional Redundancy Codes for Archival Storage
2013 IEEE 21st International Symposium on Modelling Analysis and Simulation of Computer and Telecommunication Systems, 2013Co-Authors: J F Paris, Darrell D.e. Long, Witold LitwinAbstract:Fault-tolerant disk arrays rely on replication or erasure-coding to reconstruct lost data after a disk failure. As disk capacity increases, so does the risk of encountering irrecoverable read errors that would prevent the full recovery of the lost data. We propose a three-dimensional erasure-coding technique that reduces that risk by guaranteeing full recovery in the presence of all triple and nearly all quadruple disk failures. Our solution performs better than existing solutions, such as sets of disk arrays using Reed-Solomon codes against triple failures in each individual array. Given its very high reliability, it is especially suited to the needs of very large data sets that must be preserved over long periods of time.
-
highly reliable two dimensional raid arrays for Archival Storage
International Performance Computing and Communications Conference, 2012Co-Authors: J F Paris, S Thomas J Schwarz, Ahmed Amer, Darrell D.e. LongAbstract:We present a two-dimensional RAID architecture that is specifically tailored to the needs of Archival Storage systems. Our proposal starts with a fairly conventional two-dimensional RAID architecture where each disk belongs to exactly one horizontal and one vertical RAID level 4 stripe. Once the array has been populated, we add a superparity device that contains the exclusive OR of all the contents of all horizontal-or vertical-parity disks. The new organization tolerates all triple disk failures and nearly all quadruple and quintuple disk failures. As a result, it provides mean times to data loss (MTTDLs) more than a hundred times better than those of sets of RAID level 6 stripes with equal capacity and similar parity overhead.
-
understanding data survivability in Archival Storage systems
ACM International Conference on Systems and Storage, 2012Co-Authors: Ethan L. Miller, Darrell D.e. LongAbstract:Preserving data for a long period of time in the face of faults, large and small, is crucial for designing reliable Archival Storage systems. However, the survivability of data is different from the reliability of Storage because typically, data are stored in more than one Storage at a given moment. Previous studies of reliability ignore the former. We present a framework for relating data survivability and Storage reliability, and use the framework to gauge the impact of rare but large-scale events on data survivability. We also present a method to track all copies of data and the condition of all the online and offline media, devices and systems on which they are stored uninterruptedly over the whole lifetime of the data. With this method, the survivability of the data can be closely monitored, and potential dangers can be handled in a timely manner. A better understanding of data survivability can be used in reducing unnecessary data replicas, thus reducing the cost.
-
SYSTOR - Understanding data survivability in Archival Storage systems
Proceedings of the 5th Annual International Systems and Storage Conference on - SYSTOR '12, 2012Co-Authors: Ethan L. Miller, Darrell D.e. LongAbstract:Preserving data for a long period of time in the face of faults, large and small, is crucial for designing reliable Archival Storage systems. However, the survivability of data is different from the reliability of Storage because typically, data are stored in more than one Storage at a given moment. Previous studies of reliability ignore the former. We present a framework for relating data survivability and Storage reliability, and use the framework to gauge the impact of rare but large-scale events on data survivability. We also present a method to track all copies of data and the condition of all the online and offline media, devices and systems on which they are stored uninterruptedly over the whole lifetime of the data. With this method, the survivability of the data can be closely monitored, and potential dangers can be handled in a timely manner. A better understanding of data survivability can be used in reducing unnecessary data replicas, thus reducing the cost.
Ethan L. Miller - One of the best experts on this subject based on the ideXlab platform.
-
understanding data survivability in Archival Storage systems
ACM International Conference on Systems and Storage, 2012Co-Authors: Ethan L. Miller, Darrell D.e. LongAbstract:Preserving data for a long period of time in the face of faults, large and small, is crucial for designing reliable Archival Storage systems. However, the survivability of data is different from the reliability of Storage because typically, data are stored in more than one Storage at a given moment. Previous studies of reliability ignore the former. We present a framework for relating data survivability and Storage reliability, and use the framework to gauge the impact of rare but large-scale events on data survivability. We also present a method to track all copies of data and the condition of all the online and offline media, devices and systems on which they are stored uninterruptedly over the whole lifetime of the data. With this method, the survivability of the data can be closely monitored, and potential dangers can be handled in a timely manner. A better understanding of data survivability can be used in reducing unnecessary data replicas, thus reducing the cost.
-
analysis of workload behavior in scientific and historical long term data repositories
ACM Transactions on Storage, 2012Co-Authors: Ian F. Adams, Mark W. Storer, Ethan L. MillerAbstract:The scope of Archival systems is expanding beyond cheap tertiary Storage: scientific and medical data is increasingly digital, and the public has a growing desire to digitally record their personal histories. Driven by the increase in cost efficiency of hard drives, and the rise of the Internet, content archives have become a means of providing the public with fast, cheap access to long-term data. Unfortunately, designers of purpose-built Archival systems are either forced to rely on workload behavior obtained from a narrow, anachronistic view of archives as simply cheap tertiary Storage, or extrapolate from marginally related enterprise workload data and traditional library access patterns. To close this knowledge gap and provide relevant input for the design of effective long-term data Storage systems, we studied the workload behavior of several systems within this expanded Archival Storage space. Our study examined several scientific and historical archives, covering a mixture of purposes, media types, and access models---that is, public versus private. Our findings show that, for more traditional private scientific Archival Storage, files have become larger, but update rates have remained largely unchanged. However, in the public content archives we observed, we saw behavior that diverges from the traditional “write-once, read-maybe” behavior of tertiary Storage. Our study shows that the majority of such data is modified---sometimes unnecessarily---relatively frequently, and that indexing services such as Google and internal data management processes may routinely access large portions of an archive, accounting for most of the accesses. Based on these observations, we identify areas for improving the efficiency and performance of Archival Storage systems.
-
SYSTOR - Understanding data survivability in Archival Storage systems
Proceedings of the 5th Annual International Systems and Storage Conference on - SYSTOR '12, 2012Co-Authors: Ethan L. Miller, Darrell D.e. LongAbstract:Preserving data for a long period of time in the face of faults, large and small, is crucial for designing reliable Archival Storage systems. However, the survivability of data is different from the reliability of Storage because typically, data are stored in more than one Storage at a given moment. Previous studies of reliability ignore the former. We present a framework for relating data survivability and Storage reliability, and use the framework to gauge the impact of rare but large-scale events on data survivability. We also present a method to track all copies of data and the condition of all the online and offline media, devices and systems on which they are stored uninterruptedly over the whole lifetime of the data. With this method, the survivability of the data can be closely monitored, and potential dangers can be handled in a timely manner. A better understanding of data survivability can be used in reducing unnecessary data replicas, thus reducing the cost.
-
semantic data placement for power management in Archival Storage
Petascale Data Storage Workshop, 2010Co-Authors: Avani Wildani, Ethan L. MillerAbstract:Power is the greatest lifetime cost in an Archival system, and, as decreasing costs make disks more attractive than tapes, spinning disks account for the majority of power drawn. To reduce this cost, we propose reducing the number of times disks have to spin up by grouping together files such that a typical spin-up handles several file accesses. For a typical system, we show that if only 30% of total accesses occur while disks are still spinning, we can conserve 12% of the power cost. We classify files according to directory structure and see access hit rates of up to 66% for a power savings of up to 52% of the power cost of spinning up for every read in easily-separable workloads.
-
examining energy use in heterogeneous Archival Storage systems
Modeling Analysis and Simulation On Computer and Telecommunication Systems, 2010Co-Authors: Ian F. Adams, Ethan L. Miller, Mark W. StorerAbstract:Controlling energy usage in data centers, and Storage in particular, continues to rise in importance. Many systems and models have examined energy efficiency through intelligent spin-down of disks and novel data layouts, yet little work has been done to examine how power usage over the course of months to years is impacted by the characteristics of the Storage devices chosen for use. Long-term power usage is particularly important for Archival Storage systems, since it is a large contributor to overall system cost. In this work, we begin exploring the impact that broad policies (e.g. utilize high-bandwidth devices first) have upon the power efficiency of a disk based Archival Storage system of heterogeneous devices over the course of a year. Using a discrete event simulator, we found that even simple heuristic policies for allocating space can have significant impact on the power usage of a system. We show that our system growth policies can cause power usage to vary from 10% higher to 18% lower than a naive random data allocation scheme. We also found that under low read rates power is dominated by that used in standby modes. Most interestingly, we found cases where concentrating data on fewer devices yielded increased power usage.
Mark W. Storer - One of the best experts on this subject based on the ideXlab platform.
-
analysis of workload behavior in scientific and historical long term data repositories
ACM Transactions on Storage, 2012Co-Authors: Ian F. Adams, Mark W. Storer, Ethan L. MillerAbstract:The scope of Archival systems is expanding beyond cheap tertiary Storage: scientific and medical data is increasingly digital, and the public has a growing desire to digitally record their personal histories. Driven by the increase in cost efficiency of hard drives, and the rise of the Internet, content archives have become a means of providing the public with fast, cheap access to long-term data. Unfortunately, designers of purpose-built Archival systems are either forced to rely on workload behavior obtained from a narrow, anachronistic view of archives as simply cheap tertiary Storage, or extrapolate from marginally related enterprise workload data and traditional library access patterns. To close this knowledge gap and provide relevant input for the design of effective long-term data Storage systems, we studied the workload behavior of several systems within this expanded Archival Storage space. Our study examined several scientific and historical archives, covering a mixture of purposes, media types, and access models---that is, public versus private. Our findings show that, for more traditional private scientific Archival Storage, files have become larger, but update rates have remained largely unchanged. However, in the public content archives we observed, we saw behavior that diverges from the traditional “write-once, read-maybe” behavior of tertiary Storage. Our study shows that the majority of such data is modified---sometimes unnecessarily---relatively frequently, and that indexing services such as Google and internal data management processes may routinely access large portions of an archive, accounting for most of the accesses. Based on these observations, we identify areas for improving the efficiency and performance of Archival Storage systems.
-
examining energy use in heterogeneous Archival Storage systems
Modeling Analysis and Simulation On Computer and Telecommunication Systems, 2010Co-Authors: Ian F. Adams, Ethan L. Miller, Mark W. StorerAbstract:Controlling energy usage in data centers, and Storage in particular, continues to rise in importance. Many systems and models have examined energy efficiency through intelligent spin-down of disks and novel data layouts, yet little work has been done to examine how power usage over the course of months to years is impacted by the characteristics of the Storage devices chosen for use. Long-term power usage is particularly important for Archival Storage systems, since it is a large contributor to overall system cost. In this work, we begin exploring the impact that broad policies (e.g. utilize high-bandwidth devices first) have upon the power efficiency of a disk based Archival Storage system of heterogeneous devices over the course of a year. Using a discrete event simulator, we found that even simple heuristic policies for allocating space can have significant impact on the power usage of a system. We show that our system growth policies can cause power usage to vary from 10% higher to 18% lower than a naive random data allocation scheme. We also found that under low read rates power is dominated by that used in standby modes. Most interestingly, we found cases where concentrating data on fewer devices yielded increased power usage.
-
MASCOTS - Examining Energy Use in Heterogeneous Archival Storage Systems
2010 IEEE International Symposium on Modeling Analysis and Simulation of Computer and Telecommunication Systems, 2010Co-Authors: Ian F. Adams, Ethan L. Miller, Mark W. StorerAbstract:Controlling energy usage in data centers, and Storage in particular, continues to rise in importance. Many systems and models have examined energy efficiency through intelligent spin-down of disks and novel data layouts, yet little work has been done to examine how power usage over the course of months to years is impacted by the characteristics of the Storage devices chosen for use. Long-term power usage is particularly important for Archival Storage systems, since it is a large contributor to overall system cost. In this work, we begin exploring the impact that broad policies (e.g. utilize high-bandwidth devices first) have upon the power efficiency of a disk based Archival Storage system of heterogeneous devices over the course of a year. Using a discrete event simulator, we found that even simple heuristic policies for allocating space can have significant impact on the power usage of a system. We show that our system growth policies can cause power usage to vary from 10% higher to 18% lower than a naive random data allocation scheme. We also found that under low read rates power is dominated by that used in standby modes. Most interestingly, we found cases where concentrating data on fewer devices yielded increased power usage.
-
potshards a secure recoverable long term Archival Storage system
ACM Transactions on Storage, 2009Co-Authors: Mark W. Storer, Kevin M. Greenan, Ethan L. Miller, Kaladhar VorugantiAbstract:Users are storing ever-increasing amounts of information digitally, driven by many factors including government regulations and the public's desire to digitally record their personal histories. Unfortunately, many of the security mechanisms that modern systems rely upon, such as encryption, are poorly suited for storing data for indefinitely long periods of time; it is very difficult to manage keys and update cryptosystems to provide secrecy through encryption over periods of decades. Worse, an adversary who can compromise an archive need only wait for cryptanalysis techniques to catch up to the encryption algorithm used at the time of the compromise in order to obtain “secure” data. To address these concerns, we have developed POTSHARDS, an Archival Storage system that provides long-term security for data with very long lifetimes without using encryption. Secrecy is achieved by using unconditionally secure secret splitting and spreading the resulting shares across separately managed archives. Providing availability and data recovery in such a system can be difficult; thus, we use a new technique, approximate pointers, in conjunction with secure distributed RAID techniques to provide availability and reliability across independent archives. To validate our design, we developed a prototype POTSHARDS implementation. In addition to providing us with an experimental testbed, this prototype helped us to understand the design issues that must be addressed in order to maximize security.
-
Secure, energy-efficient, evolvable, long-term Archival Storage
2009Co-Authors: Ethan L. Miller, Mark W. StorerAbstract:Users are storing ever-increasing amounts of information digitally, driven by many factors including government regulations and the public's desire to digitally record their personal histories. Unfortunately, we have yet to demonstrate that we can reliably preserve digital data for more than a few years, putting a generation's cultural legacy at risk. Much of the problem is rooted in our approach to building long-term Storage systems; currently Archival systems are developed using the same approaches, access patterns and techniques used to design higher-performance, shorter-term Storage systems. As a result, current Archival Storage systems still rely on strategies that fail in long-term scenarios, waste money and energy, and perpetuate the endless cycles of “fork-lift” upgrades and wholesale migrations needed to remain efficient and up to date. In my thesis, I demonstrate that Archival Storage is a first class category of Storage that requires specialized solutions. To this end, I present several techniques tailored specifically for the unique demands of long-lived data. To explore the security needs of Archival data, I have developed POTSHARDS, which offers secrecy through unconditionally secure secrecy techniques, and survivability through increased attack detection and built-in data recovery. To study cost savings, I have created Pergamum, a distributed system of intelligent Storage appliances that stores data reliably with multi-level encoding and a hierarchical auditing scheme, and energy-efficiently by leveraging existing MAID techniques, while extending them by exploiting the different access patterns of data and metadata. Running atop of Pergamum is Logan, a management layer being developed that actively identifies and decommissions wasteful devices in order to continuously maximize system efficiency. These systems combine to demonstrate significant progress towards effective, secure, energy-efficient, and evolvable Archival Storage.
K Gopinath - One of the best experts on this subject based on the ideXlab platform.
-
g_ its 2 vsr an information theoretical secure verifiable secret redistribution protocol for long term Archival Storage
Fourth International IEEE Security in Storage Workshop, 2007Co-Authors: V H Gupta, K GopinathAbstract:Protocols for secure Archival Storage are becoming increasingly important as the use of digital Storage for sensitive documents is gaining wider practice. In [8], Wong et al. combined verifiable secret sharing with proactive secret sharing without reconstruction and proposed a verifiable secret redistribution protocol for long term Storage. However, their protocol requires that each of the receivers is honest during redistribution. We proposed [3] an extension to their protocol wherein we relaxed the requirement that all the recipients should be honest to the condition that only a simple majority amongst the recipients need to be honest during the re(distribution) processes. Further, both of these protocols make use of Feldman 's approach for achieving integrity during the (re)distribution processes. In this paper, we present a revised version of our earlier protocol, and its adaptation to incorporate Pedersen 's approach instead of Feldman's thereby achieving information theoretic secrecy while retaining integrity guarantees.
-
IEEE Security in Storage Workshop - G_{its}^2 VSR: An Information Theoretical Secure Verifiable Secret Redistribution Protocol for Long-term Archival Storage
Fourth International IEEE Security in Storage Workshop, 2007Co-Authors: V H Gupta, K GopinathAbstract:Protocols for secure Archival Storage are becoming increasingly important as the use of digital Storage for sensitive documents is gaining wider practice. In [8], Wong et al. combined verifiable secret sharing with proactive secret sharing without reconstruction and proposed a verifiable secret redistribution protocol for long term Storage. However, their protocol requires that each of the receivers is honest during redistribution. We proposed [3] an extension to their protocol wherein we relaxed the requirement that all the recipients should be honest to the condition that only a simple majority amongst the recipients need to be honest during the re(distribution) processes. Further, both of these protocols make use of Feldman 's approach for achieving integrity during the (re)distribution processes. In this paper, we present a revised version of our earlier protocol, and its adaptation to incorporate Pedersen 's approach instead of Feldman's thereby achieving information theoretic secrecy while retaining integrity guarantees.
V H Gupta - One of the best experts on this subject based on the ideXlab platform.
-
g_ its 2 vsr an information theoretical secure verifiable secret redistribution protocol for long term Archival Storage
Fourth International IEEE Security in Storage Workshop, 2007Co-Authors: V H Gupta, K GopinathAbstract:Protocols for secure Archival Storage are becoming increasingly important as the use of digital Storage for sensitive documents is gaining wider practice. In [8], Wong et al. combined verifiable secret sharing with proactive secret sharing without reconstruction and proposed a verifiable secret redistribution protocol for long term Storage. However, their protocol requires that each of the receivers is honest during redistribution. We proposed [3] an extension to their protocol wherein we relaxed the requirement that all the recipients should be honest to the condition that only a simple majority amongst the recipients need to be honest during the re(distribution) processes. Further, both of these protocols make use of Feldman 's approach for achieving integrity during the (re)distribution processes. In this paper, we present a revised version of our earlier protocol, and its adaptation to incorporate Pedersen 's approach instead of Feldman's thereby achieving information theoretic secrecy while retaining integrity guarantees.
-
IEEE Security in Storage Workshop - G_{its}^2 VSR: An Information Theoretical Secure Verifiable Secret Redistribution Protocol for Long-term Archival Storage
Fourth International IEEE Security in Storage Workshop, 2007Co-Authors: V H Gupta, K GopinathAbstract:Protocols for secure Archival Storage are becoming increasingly important as the use of digital Storage for sensitive documents is gaining wider practice. In [8], Wong et al. combined verifiable secret sharing with proactive secret sharing without reconstruction and proposed a verifiable secret redistribution protocol for long term Storage. However, their protocol requires that each of the receivers is honest during redistribution. We proposed [3] an extension to their protocol wherein we relaxed the requirement that all the recipients should be honest to the condition that only a simple majority amongst the recipients need to be honest during the re(distribution) processes. Further, both of these protocols make use of Feldman 's approach for achieving integrity during the (re)distribution processes. In this paper, we present a revised version of our earlier protocol, and its adaptation to incorporate Pedersen 's approach instead of Feldman's thereby achieving information theoretic secrecy while retaining integrity guarantees.