The Experts below are selected from a list of 10197 Experts worldwide ranked by ideXlab platform

Hani Hamdan - One of the best experts on this subject based on the ideXlab platform.

  • model selection with bic and icl criteria for Binned Data clustering by bin em cem algorithms
    Systems Man and Cybernetics, 2013
    Co-Authors: Hani Hamdan
    Abstract:

    Several clustering approaches are adapted to Binned Data in order to accelerate the clustering process or to deal with Data of limited precision. Bin-EM-CEM algorithms of fourteen parsimonious Gaussian mixture models are developed. Each model performs differently according to its specific feature. Without knowing any information of the Data, a criterion is considered to select the best model in order to obtain a good result. In this article, BIC and ICL criteria are adapted to Binned Data clustering to choose the bin-EM-CEM algorithm of the right model as well as the number of clusters. By different experiments on simulated Data and real Data, the performance of BIC and ICL criteria in model selection for Binned Data clustering are studied and compared on different aspects.

  • Bin-EM-CEM algorithms of general parsimonious Gaussian mixture models for Binned Data clustering
    2013
    Co-Authors: Hani Hamdan
    Abstract:

    Data binning is a well-known Data pre-processing technique in statistics. It was applied to model-based clustering approaches to reduce the number of Data and facilitate the processing. EM and CEM algorithms are commonly used in model-based approaches. Thus EM and CEM algorithms applied to Binned Data were developed: Binned-EM algorithm for mixture approach, and bin-EM-CEM algorithm for classification approach. At another side, fourteen parsimonious Gaussian mixture models for EM and CEM algorithms were proposed by considering a parametrization of the variance matrices of the clusters. Due to different characteristics of each model, fourteen models can adapt to Data of different structures so as to simplify the clustering process. The experimental results of EM algorithms of fourteen parsimonious models also show that the model which fits the Data gives a better result than the other models. Previously, Binned-EM algorithms of fourteen parsimonious Gaussian mixture models were developed. The result shows to be of interest to combine the advantages of Binned Data and parsimonious models on model-based clustering approaches. So in this paper, we develop bin-EM-CEM algorithms of the eight most general parsimonious Gaussian mixture models. The performances of the developed algorithms applied to different models of Data are studied and analyzed.

  • Bin-EM-CEM algorithms of spherical parsimonious Gaussian mixture models for Binned Data clustering
    2013
    Co-Authors: Hani Hamdan
    Abstract:

    EM algorithm is widely used in clustering domain because of its easy implementation and small storage space. CEM algorithm, which is considered as a classification version of EM algorithm, is another common used clustering algorithm. With the development of technology, we obtain more and more Data. This results in slow computation of EM and CEM algorithms. Binning Data seems to be efficient in gaining computation time by reducing the number of observations to the number of bins. Thus, EM and CEM algorithms applied to Binned Data were proposed: Binned-EM and bin-EM-CEM algorithms. Moreover, fourteen parsimonious Gaussian mixture models, generated according to eigenvalue decomposition of the variance matrices of the mixture components, have less parameters than the most general model. By applying the EM and CEM algorithms of parsimonious models, estimation process is simplified and then accelerated. In this paper, to combine the advantages of Binned Data and parsimonious Gaussian mixture models, we develop bin-EM-CEM algorithms of spherical parsimonious Gaussian mixture models.

  • mixture model clustering of Binned uncertain Data
    Journal of Control Engineering and Applied Informatics, 2012
    Co-Authors: Hani Hamdan
    Abstract:

    This paper addresses the problem of taking into account Data imprecision in the mixture model clustering of Binned Data. Binning (or grouping) Data is common in Data analysis and machine learning. Recently, we developed an original method which fitted the binning Data procedure to imprecise Data. The idea was to model imprecise Data by multivariate uncertainty zones and to assign each uncertainty zone to several bins with proportions proportional to its overlapping volumes with the bins. The experimental results of this method when it was associated with the Binned-EM algorithm (mixture approach) were encouraging. However, the Binned-EM algorithm has the disadvantage of being sometimes computationally expensive. To overcome this problem, we propose in this paper to apply our binning Data procedure with the classification approach based on bin-EM-CEM algorithm which is much faster than the Binned-EM algorithm. The paper concludes with a brief description of a flaw diagnosis application using acoustic emission. The experimental results compare our binning Data procedure with the classical one (when applied to imprecise Data) in the classification approach framework, and with the int-EM-CEM algorithm, in the context of Binned bivariate measurements of acoustic emission event localization.

  • Parsimonious Gaussian mixture models of general family for Binned Data clustering: mixture approach
    2012
    Co-Authors: Hani Hamdan
    Abstract:

    Binning Data provides a solution in deducing computation expense in cluster analysis. According to former study, basing cluster analysis on Gaussian mixture models has become a classical and power approach. Mixture approach is one of the most common model-based approaches, which estimates the model parameters by maximizing the likelihood by EM algorithm. According to eigenvalue composition of the variance matrices of the mixture components, parsimonious models are generated. Choosing a right parsimonious model is crucial in obtaining a good result. In this paper, we address the problem of applying mixture approach to Binned Data (Binned-EM algorithm). Six general models are studied and the difference in the performances of six general models is analyzed.

Stasa Milojevic - One of the best experts on this subject based on the ideXlab platform.

  • power law distributions in information science making the case for logarithmic binning
    Journal of the Association for Information Science and Technology, 2010
    Co-Authors: Stasa Milojevic
    Abstract:

    We suggest partial logarithmic binning as the method of choice for uncovering the nature of many distributions encountered in information science (IS). Logarithmic binning retrieves information and trends “not visible” in noisy power law tails. We also argue that obtaining the exponent from logarithmically Binned Data using a simple least square method is in some cases warranted in addition to methods such as the maximum likelihood. We also show why often-used cumulative distributions can make it difficult to distinguish noise from genuine features and to obtain an accurate power law exponent of the underlying distribution. The treatment is nontechnical, aimed at IS researchers with little or no background in mathematics. © 2010 Wiley Periodicals, Inc.

  • power law distributions in information science making the case for logarithmic binning
    arXiv: Physics and Society, 2010
    Co-Authors: Stasa Milojevic
    Abstract:

    We suggest partial logarithmic binning as the method of choice for uncovering the nature of many distributions encountered in information science (IS). Logarithmic binning retrieves information and trends "not visible" in noisy power-law tails. We also argue that obtaining the exponent from logarithmically Binned Data using a simple least square method is in some cases warranted in addition to methods such as the maximum likelihood. We also show why often used cumulative distributions can make it difficult to distinguish noise from genuine features, and make it difficult to obtain an accurate power-law exponent of the underlying distribution. The treatment is non-technical, aimed at IS researchers with little or no background in mathematics.

Tong-jie Zhang - One of the best experts on this subject based on the ideXlab platform.

  • cosmological model independent test of varlambda cdm with two point diagnostic by the observational hubble parameter Data
    European Physical Journal C, 2018
    Co-Authors: Shu-lei Cao, Xiao-wei Duan, Xiao-lei Meng, Tong-jie Zhang
    Abstract:

    Aiming at exploring the nature of dark energy (DE), we use forty-three observational Hubble parameter Data (OHD) in the redshift range $$0 < z \leqslant 2.36$$ to make a cosmological model-independent test of the $$\varLambda $$ CDM model with two-point $$Omh^2(z_{2};z_{1})$$ diagnostic. In $$\varLambda $$ CDM model, with equation of state (EoS) $$w=-1$$ , two-point diagnostic relation $$Omh^2 \equiv \varOmega _{\mathrm{m}}h^2$$ is tenable, where $$\varOmega _{\mathrm{m}}$$ is the present matter density parameter, and h is the Hubble parameter divided by 100 $${\mathrm {km\, s^{-1} \ Mpc^{-1}}}$$ . We utilize two methods: the weighted mean and median statistics to bin the OHD to increase the signal-to-noise ratio of the measurements. The binning methods turn out to be promising and considered to be robust. By applying the two-point diagnostic to the Binned Data, we find that although the best-fit values of $$Omh^2$$ fluctuate as the continuous redshift intervals change, on average, they are continuous with being constant within 1 $$\sigma $$ confidence interval. Therefore, we conclude that the $$\varLambda $$ CDM model cannot be ruled out.

  • cosmological model independent test of lambda cdm with two point diagnostic by the observational hubble parameter Data
    arXiv: Cosmology and Nongalactic Astrophysics, 2017
    Co-Authors: Shu-lei Cao, Xiao-wei Duan, Xiao-lei Meng, Tong-jie Zhang
    Abstract:

    Aiming at exploring the nature of dark energy (DE), we use forty-three observational Hubble parameter Data (OHD) in the redshift range $0 < z \leqslant 2.36$ to make a cosmological model-independent test of the $\Lambda$CDM model with two-point $Omh^2(z_{2};z_{1})$ diagnostic. In $\Lambda$CDM model, with equation of state (EoS) $w=-1$, two-point diagnostic relation $Omh^2 \equiv \Omega_m h^2$ is tenable, where $\Omega_m$ is the present matter density parameter, and $h$ is the Hubble parameter divided by 100 $\rm km s^{-1} Mpc^{-1}$. We utilize two methods: the weighted mean and median statistics to bin the OHD to increase the signal-to-noise ratio of the measurements. The binning methods turn out to be promising and considered to be robust. By applying the two-point diagnostic to the Binned Data, we find that although the best-fit values of $Omh^2$ fluctuate as the continuous redshift intervals change, on average, they are continuous with being constant within 1 $\sigma$ confidence interval. Therefore, we conclude that the $\Lambda$CDM model cannot be ruled out.

Gerard Govaert - One of the best experts on this subject based on the ideXlab platform.

  • a classification em algorithm for Binned Data
    Computational Statistics & Data Analysis, 2006
    Co-Authors: Allou Same, Christophe Ambroise, Gerard Govaert
    Abstract:

    A real-time flaw diagnosis application for pressurized containers using acoustic emissions is described. The pressurized containers used are cylindrical tanks containing fluids under pressure. The surface of the pressurized containers is divided into bins, and the number of acoustic signals emanating from each bin is counted. Spatial clustering of high density bins using mixture models is used to detect flaws. A dedicated EM algorithm can be derived to select the mixture parameters, but this is a greedy algorithm since it requires the numerical computation of integrals and may converge only slowly. To deal with this problem, a classification version of the EM (CEM) algorithm is defined, and using synthetic and real Data sets, the proposed algorithm is compared to the CEM algorithm applied to classical Data. The two approaches generate comparable solutions in terms of the resulting partition if the histogram is sufficiently accurate, but the algorithm designed for Binned Data becomes faster when the number of available observations is large enough.

  • the fitting of Binned Data clustering to imprecise Data
    International Conference on Information and Communication Technologies, 2004
    Co-Authors: Hani Hamdan, Gerard Govaert
    Abstract:

    This paper addresses the problem of taking into account the Data imprecision in the clustering of Binned Data using mixture models and Binned EM algorithm. Within the framework of a defects detection problem by acoustic emission control, we were brought to treat a set of points using the EM algorithm applied to a diagonal Gaussian mixture model. This one provides a satisfactory solution but the real time constraints imposed in our problem make its application impossible when the number of points becomes too big. As Data sets become larger, Data processing becomes increasingly complex and as a result, the Data analysis is expensive in computation time. The solution that we propose is to group Data and available Data thus takes the form of a histogram. Such Data are also called Binned Data. We fit the binning Data procedure to imprecise Data. We model imprecise Data by multivariate uncertainty zones and we propose to assign each uncertainty zone to several bins with percentages proportional to its overlapping surfaces with the bins. The experimental results compare this binning procedure with the classical one (applied to imprecise points) and with the interval EM algorithm considered here as a reference, using simulated Data.

  • a mixture model approach for Binned Data clustering
    Intelligent Data Analysis, 2003
    Co-Authors: Allou Same, Christophe Ambroise, Gerard Govaert
    Abstract:

    In some particular Data analysis problems, available Data takes the form of an histogram. Such Data are also called Binned Data. This paper addresses the problem of clustering Binned Data using mixture models. A specific EM algorithm has been proposed by Cadez et al.([2]) to deal with these Data. This algorithm has the disadvantage of being computationally expensive. In this paper, a classification version of this algorithm is proposed, which is much faster. The two approaches are compared using simulated Data. The simulation results show that both algorithms generate comparable solutions in terms of resulting partition if the histogram is accurate enough.

Deng Wen-shuenn - One of the best experts on this subject based on the ideXlab platform.

  • A Note on Frequency Polygon Based on Weighted Sum of Binned Data
    'Informa UK Limited', 2018
    Co-Authors: Deng Wen-shuenn
    Abstract:

    [[abstract]]We revisit the generalized midpoint frequency polygons of Scott (1985), and the edge frequency polygons of Jones et al. (1998 Jones, M.C., Samiuddin, M., Al-Harbey, A.H., Maatouk, T. A.H. (1998). The edge frequency polygon. Biometrika 85:235–239. [Crossref], [Web of Science ®], [Google Scholar] ) and Dong and Zheng (2001 Dong, J.P., Zheng, C. (2001). Generalized edge frequency polygon for density estimation. Statist. Probab. Lett. 55:137–145. [Crossref], [Web of Science ®], [Google Scholar] ). Their estimators are linear interpolants of the appropriate values above the bin centers or edges, those values being weighted averages of the heights of r, r ∈ N, neighboring histogram bins. We propose a simple kernel evaluation method to generate weights for Binned values. The proposed kernel method can provide near-optimal weights in the sense of minimizing asymptotic mean integrated square error. In addition, we prove that the discrete uniform weights minimize the variance of the generalized frequency polygon under some mild conditions. Analogous results are obtained for the generalized frequency polygon based on linearly preBinned Data. Finally, we use two examples and a simulation study to compare the generalized midpoint and edge frequency polygons.[[notice]]補正完

  • A Note on the Frequency Polygon Based on Weighted Sum of Binned Data
    'Informa UK Limited', 2016
    Co-Authors: Deng Wen-shuenn
    Abstract:

    [[abstract]]We revisit the generalized midpoint frequency polygons of Scott (1985), and the edge frequency polygons of Jones et al. (1998) and Dong and Zheng (2001). Their estimators are linear interpolants of the appropriate values above the bin centers or edges, those values being weighted averages of the heights of r, r ∈ N, neighboring histogram bins. We propose a simple kernel evaluation method to generate weights for Binned values. The proposed kernel method can provide near-optimal weights in the sense ofminimizing asymptotic mean integrated square error. In addition, we prove that the discrete uniform weights minimize the variance of the generalized frequency polygon under some mild conditions. Analogous results are obtained for the generalized frequency polygon based on linearly preBinned Data. Finally, we use two examples and a simulation study to compare the generalized midpoint and edge frequency polygons.[[notice]]補正完

  • A Note on the Frequency Polygon Based on the Weighted Sums of Binned Data
    'Informa UK Limited', 2014
    Co-Authors: Deng Wen-shuenn
    Abstract:

    [[abstract]]We revisit the generalized midpoint frequency polygons of Scott (1985), and the edge frequency polygons of Jones et al. (1998) and Dong and Zheng (2001). Their estimators are linear interpolants of the appropriate values above the bin centers or edges, those values being weighted averages of the heights of r, r ∈ N, neighboring histogram bins. We propose a simple kernel evaluation method to generate weights for Binned values. The proposed kernel method can provide near-optimal weights in the sense ofminimizing asymptotic mean integrated square error. In addition, we prove that the discrete uniform weights minimize the variance of the generalized frequency polygon under some mild conditions. Analogous results are obtained for the generalized frequency polygon based on linearly preBinned Data. Finally, we use two examples and a simulation study to compare the generalized midpoint and edge frequency polygons.[[notice]]補正完畢[[journaltype]]國外[[incitationindex]]SSCI[[ispeerreviewed]]Y[[booktype]]紙本[[countrycodes]]US