The Experts below are selected from a list of 16287 Experts worldwide ranked by ideXlab platform

Ruiz B Cobo - One of the best experts on this subject based on the ideXlab platform.

  • an open source Massively parallel code for non lte synthesis and inversion of spectral lines and zeeman induced stokes profiles
    Astronomy and Astrophysics, 2015
    Co-Authors: H Socasnavarro, J De La Cruz Rodriguez, Asensio A Ramos, Trujillo J Bueno, Ruiz B Cobo
    Abstract:

    With the advent of a new generation of solar telescopes and instrumentation, interpreting chromospheric observations (in particular, spectropolarimetry) requires new, suitable diagnostic tools. This paper describes a new code, NICOLE, that has been designed for Stokes non-LTE radiative transfer, for synthesis and inversion of spectral lines and Zeeman-induced polarization profiles, spanning a wide range of atmospheric heights from the photosphere to the chromosphere. The code features a number of unique features and capabilities and has been built from scratch with a powerful parallelization scheme that makes it suitable for application on Massive Datasets using large supercomputers. The source code is written entirely in Fortran 90/2003 and complies strictly with the ANSI standards to ensure maximum compatibility and portability. It is being publicly released, with the idea of facilitating future branching by other groups to augment its capabilities.

  • an open source Massively parallel code for non lte synthesis and inversion of spectral lines and zeeman induced stokes profiles
    arXiv: Solar and Stellar Astrophysics, 2014
    Co-Authors: H Socasnavarro, J De La Cruz Rodriguez, Asensio A Ramos, Trujillo J Bueno, Ruiz B Cobo
    Abstract:

    With the advent of a new generation of solar telescopes and instrumentation, the interpretation of chromospheric observations (in particular, spectro-polarimetry) requires new, suitable diagnostic tools. This paper describes a new code, NICOLE, that has been designed for Stokes non-LTE radiative transfer, both for synthesis and inversion of spectral lines and Zeeman-induced polarization profiles, spanning a wide range of atmospheric heights, from the photosphere to the chromosphere. The code fosters a number of unique features and capabilities and has been built from scratch with a powerful parallelization scheme that makes it suitable for application on Massive Datasets using large supercomputers. The source code is being publicly released, with the idea of facilitating future branching by other groups to augment its capabilities.

Dennis K J Lin - One of the best experts on this subject based on the ideXlab platform.

  • Data skeletons: simultaneous estimation of multiple quantiles for Massive streaming Datasets with applications to density estimation
    Statistics and Computing, 2007
    Co-Authors: James P Mcdermott, John C. Liechty, G. Jogesh Babu, Dennis K J Lin
    Abstract:

    We consider the problem of density estimation when the data is in the form of a continuous stream with no fixed length. In this setting, implementations of the usual methods of density estimation such as kernel density estimation are problematic. We propose a method of density estimation for Massive Datasets that is based upon taking the derivative of a smooth curve that has been fit through a set of quantile estimates. To achieve this, a low-storage, single-pass, sequential method is proposed for simultaneous estimation of multiple quantiles for Massive Datasets that form the basis of this method of density estimation. For comparison, we also consider a sequential kernel density estimator. The proposed methods are shown through simulation study to perform well and to have several distinct advantages over existing methods.

  • regression analysis for Massive Datasets
    Data and Knowledge Engineering, 2007
    Co-Authors: Tsaihung Fan, Dennis K J Lin, K F Cheng
    Abstract:

    In the past decades, we have witnessed a revolution in information technology. Routine collection of systematically generated data is now commonplace. Databases with hundreds of fields (variables), and billions of records (observations) are not unusual. This presents a difficulty for classical data analysis methods, mainly due to the limitation of computer memory and computational costs (in time, for example). In this paper, we propose an intelligent regression analysis methodology which is suitable for modeling Massive Datasets. The basic idea here is to split the entire dataset into several blocks, applying the classical regression techniques for data in each block, and finally combining these regression results via weighted averages. Theoretical justification of the goodness of the proposed method is given, and empirical performance based on extensive simulation study is discussed.

  • quantile contours and multivariate density estimation for Massive Datasets via sequential convex hull peeling
    Iie Transactions, 2007
    Co-Authors: James P Mcdermott, Dennis K J Lin
    Abstract:

    We propose a low-storage, single-pass, sequential method for the execution of convex hull peeling for Massive Datasets. The method is shown to vastly reduce the computation time required for the existing convex hull peeling algorithm from O(n 2) to O(n). Furthermore, the proposed method has significantly smaller storage requirements compared to the existing method. We present algorithms for low-storage, sequential computation of both the convex hull peeling multivariate median and the convex hull peeling pth depth contour, where 0 < p < 1. We demonstrate the accuracy and reduced computation time required of the proposed method by comparing to the existing convex hull peeling method through simulation studies.

  • Single-pass low-storage arbitrary quantile estimation for Massive Datasets
    Statistics and Computing, 2003
    Co-Authors: John C. Liechty, Dennis K J Lin, James P Mcdermott
    Abstract:

    We present a single-pass, low-storage, sequential method for estimating an arbitrary quantile of an unknown distribution. The proposed method performs very well when compared to existing methods for estimating the median as well as arbitrary quantiles for a wide range of densities. In addition to explaining the method and presenting the results of the simulation study, we discuss intuition behind the method and demonstrate empirically, for certain densities, that the proposed estimator converges to the sample quantile.

James P Mcdermott - One of the best experts on this subject based on the ideXlab platform.

  • Data skeletons: simultaneous estimation of multiple quantiles for Massive streaming Datasets with applications to density estimation
    Statistics and Computing, 2007
    Co-Authors: James P Mcdermott, John C. Liechty, G. Jogesh Babu, Dennis K J Lin
    Abstract:

    We consider the problem of density estimation when the data is in the form of a continuous stream with no fixed length. In this setting, implementations of the usual methods of density estimation such as kernel density estimation are problematic. We propose a method of density estimation for Massive Datasets that is based upon taking the derivative of a smooth curve that has been fit through a set of quantile estimates. To achieve this, a low-storage, single-pass, sequential method is proposed for simultaneous estimation of multiple quantiles for Massive Datasets that form the basis of this method of density estimation. For comparison, we also consider a sequential kernel density estimator. The proposed methods are shown through simulation study to perform well and to have several distinct advantages over existing methods.

  • quantile contours and multivariate density estimation for Massive Datasets via sequential convex hull peeling
    Iie Transactions, 2007
    Co-Authors: James P Mcdermott, Dennis K J Lin
    Abstract:

    We propose a low-storage, single-pass, sequential method for the execution of convex hull peeling for Massive Datasets. The method is shown to vastly reduce the computation time required for the existing convex hull peeling algorithm from O(n 2) to O(n). Furthermore, the proposed method has significantly smaller storage requirements compared to the existing method. We present algorithms for low-storage, sequential computation of both the convex hull peeling multivariate median and the convex hull peeling pth depth contour, where 0 < p < 1. We demonstrate the accuracy and reduced computation time required of the proposed method by comparing to the existing convex hull peeling method through simulation studies.

  • Single-pass low-storage arbitrary quantile estimation for Massive Datasets
    Statistics and Computing, 2003
    Co-Authors: John C. Liechty, Dennis K J Lin, James P Mcdermott
    Abstract:

    We present a single-pass, low-storage, sequential method for estimating an arbitrary quantile of an unknown distribution. The proposed method performs very well when compared to existing methods for estimating the median as well as arbitrary quantiles for a wide range of densities. In addition to explaining the method and presenting the results of the simulation study, we discuss intuition behind the method and demonstrate empirically, for certain densities, that the proposed estimator converges to the sample quantile.

David Madigan - One of the best experts on this subject based on the ideXlab platform.

  • a one pass sequential monte carlo method for bayesian analysis of Massive Datasets
    Bayesian Analysis, 2006
    Co-Authors: Suhrid Balakrishnan, David Madigan
    Abstract:

    For Bayesian analysis of Massive data, Markov chain Monte Carlo (MCMC) techniques often prove infeasible due to computational resource con- straints. Standard MCMC methods generally require a complete scan of the dataset for each iteration. Ridgeway and Madigan (2002) and Chopin (2002b) recently presented importance sampling algorithms that combined simulations from a posterior distribution conditioned on a small portion of the dataset with a reweighting of those simulations to condition on the remainder of the dataset. While these algorithms drastically reduce the number of data accesses as compared to traditional MCMC, they still require substantially more than a single pass over the dataset. In this paper, we present \1PFS," an ecien t, one-pass algorithm. The algorithm employs a simple modication of the Ridgeway and Madigan (2002) particle ltering algorithm that replaces the MCMC based \rejuvenation" step with a more ecien t \shrinkage" kernel smoothing based step. To show proof- of-concept and to enable a direct comparison, we demonstrate 1PFS on the same examples presented in Ridgeway and Madigan (2002), namely a mixture model for Markov chains and Bayesian logistic regression. Our results indicate the proposed scheme delivers accurate parameter estimates while employing only a single pass through the data.

  • a sequential monte carlo method for bayesian analysis of Massive Datasets
    Data Mining and Knowledge Discovery, 2003
    Co-Authors: Greg Ridgeway, David Madigan
    Abstract:

    Markov chain Monte Carlo (MCMC) techniques revolutionized statistical practice in the 1990s by providing an essential toolkit for making the rigor and flexibility of Bayesian analysis computationally practical. At the same time the increasing prevalence of Massive Datasets and the expansion of the field of data mining has created the need for statistically sound methods that scale to these large problems. Except for the most trivial examples, current MCMC methods require a complete scan of the dataset for each iteration eliminating their candidacy as feasible data mining techniques. In this article we present a method for making Bayesian analysis of Massive Datasets computationally feasible. The algorithm simulates from a posterior distribution that conditions on a smaller, more manageable portion of the dataset. The remainder of the dataset may be incorporated by reweighting the initial draws using importance sampling. Computation of the importance weights requires a single scan of the remaining observations. While importance sampling increases efficiency in data access, it comes at the expense of estimation efficiency. A simple modification, based on the “rejuvenation” step used in particle filters for dynamic systems models, sidesteps the loss of efficiency with only a slight increase in the number of data accesses. To show proof-of-concept, we demonstrate the method on two examples. The first is a mixture of transition models that has been used to model web traffic and robotics. For this example we show that estimation efficiency is not affected while offering a 99% reduction in data accesses. The second example applies the method to Bayesian logistic regression and yields a 98% reduction in data accesses.

  • bayesian analysis of Massive Datasets via particle filters
    Knowledge Discovery and Data Mining, 2002
    Co-Authors: Greg Ridgeway, David Madigan
    Abstract:

    Markov Chain Monte Carlo (MCMC) techniques revolutionized statistical practice in the 1990s by providing an essential toolkit for making the rigor and flexibility of Bayesian analysis computationally practical. At the same time the increasing prevalence of Massive Datasets and the expansion of the field of data mining has created the need to produce statistically sound methods that scale to these large problems. Except for the most trivial examples, current MCMC methods require a complete scan of the dataset for each iteration eliminating their candidacy as feasible data mining techniques.In this article we present a method for making Bayesian analysis of Massive Datasets computationally feasible. The algorithm simulates from a posterior distribution that conditions on a smaller, more manageable portion of the dataset. The remainder of the dataset may be incorporated by reweighting the initial draws using importance sampling. Computation of the importance weights requires a single scan of the remaining observations. While importance sampling increases efficiency in data access, it comes at the expense of estimation efficiency. A simple modification, based on the "rejuvenation" step used in particle filters for dynamic systems models, sidesteps the loss of efficiency with only a slight increase in the number of data accesses.To show proof-of-concept, we demonstrate the method on a mixture of transition models that has been used to model web traffic and robotics. For this example we show that estimation efficiency is not affected while offering a 95% reduction in data accesses.

H Socasnavarro - One of the best experts on this subject based on the ideXlab platform.

  • an open source Massively parallel code for non lte synthesis and inversion of spectral lines and zeeman induced stokes profiles
    Astronomy and Astrophysics, 2015
    Co-Authors: H Socasnavarro, J De La Cruz Rodriguez, Asensio A Ramos, Trujillo J Bueno, Ruiz B Cobo
    Abstract:

    With the advent of a new generation of solar telescopes and instrumentation, interpreting chromospheric observations (in particular, spectropolarimetry) requires new, suitable diagnostic tools. This paper describes a new code, NICOLE, that has been designed for Stokes non-LTE radiative transfer, for synthesis and inversion of spectral lines and Zeeman-induced polarization profiles, spanning a wide range of atmospheric heights from the photosphere to the chromosphere. The code features a number of unique features and capabilities and has been built from scratch with a powerful parallelization scheme that makes it suitable for application on Massive Datasets using large supercomputers. The source code is written entirely in Fortran 90/2003 and complies strictly with the ANSI standards to ensure maximum compatibility and portability. It is being publicly released, with the idea of facilitating future branching by other groups to augment its capabilities.

  • an open source Massively parallel code for non lte synthesis and inversion of spectral lines and zeeman induced stokes profiles
    arXiv: Solar and Stellar Astrophysics, 2014
    Co-Authors: H Socasnavarro, J De La Cruz Rodriguez, Asensio A Ramos, Trujillo J Bueno, Ruiz B Cobo
    Abstract:

    With the advent of a new generation of solar telescopes and instrumentation, the interpretation of chromospheric observations (in particular, spectro-polarimetry) requires new, suitable diagnostic tools. This paper describes a new code, NICOLE, that has been designed for Stokes non-LTE radiative transfer, both for synthesis and inversion of spectral lines and Zeeman-induced polarization profiles, spanning a wide range of atmospheric heights, from the photosphere to the chromosphere. The code fosters a number of unique features and capabilities and has been built from scratch with a powerful parallelization scheme that makes it suitable for application on Massive Datasets using large supercomputers. The source code is being publicly released, with the idea of facilitating future branching by other groups to augment its capabilities.