The Experts below are selected from a list of 10284 Experts worldwide ranked by ideXlab platform

Andrei Achimaş Cadariu - One of the best experts on this subject based on the ideXlab platform.

  • Binomial distribution sample confidence intervals estimation 2 proportion like medical key parameters
    2003
    Co-Authors: Andrei Achimaş Cadariu
    Abstract:

    The accuracy of confidence interval expressing is one of most important problems in medical statistics. In the paper was considers for accuracy analysis from literature, a number of confidence interval formulas for Binomial proportions from the most used ones. For each method, based on theirs formulas, were implements in PHP an algorithm of computing. The performance of each method for different sample sizes and different values of Binomial Variable was compares using a set of criteria: the average of experimental errors, the standard deviation of errors, and the deviation relative to imposed significance level. A new method, named Binomial, based on Bayes and Jeffreys methods, was defines, implements, and tested. The results show that there is no method from considered ones, which to have a constant behavior in terms of average of the experimental errors, the standard deviation of errors, or the deviation relative to imposed significance level for every sample sizes, Binomial Variables, or proportions of them. More, this unfortunately behavior remains also for large and small subsets of natural numbers set. The Binomial method obtained systematically lowest deviations relative to imposed significance level (α = 5%) starting with a certain sample size.

  • Binomial distribution sample confidence intervals estimation 4 post test probability
    2003
    Co-Authors: Andrei Achimaş Cadariu
    Abstract:

    Posttest probability is one of the key parameters which can be measured and interpret in dichotomous diagnostic tests. Defined as the proportion of patients, which present particular test result and the target disorder, the posttest probability is a parameter used in assessing the efficiency of the diagnostic. As a point estimated parameter, posttest probability needs a confidence interval in order to interpreting trustworthiness or robustness of the finding. Unfortunately, for post test probability there was no confidence intervals reported in literature. The aim of this paper is to introduce six methods named Wilson, Logit, LogitC, BayesF, Jeffreys, and Binomial as methods of computing confidence intervals for posttest probability and to present theirs performances. Computer implementations of the methods use the PHP language. The performance of each method for different sample sizes and different values of Binomial Variable was asses using a set of criterions. One criterion was the average of experimental errors and standard deviations. Second, the deviation relative to imposed significance level (α = 5%). Third, the behavior of the methods when the sample size vary from 4 to 103 and on random sample and random Binomial Variable in 4..1000 domain. The results of the experiments show us that the Binomial method obtain the

Xinjia Chen - One of the best experts on this subject based on the ideXlab platform.

  • coverage probability of random intervals
    arXiv: Statistics Theory, 2007
    Co-Authors: Xinjia Chen
    Abstract:

    In this paper, we develop a general theory on the coverage probability of random intervals defined in terms of discrete random Variables with continuous parameter spaces. The theory shows that the minimum coverage probabilities of random intervals with respect to corresponding parameters are achieved at discrete finite sets and that the coverage probabilities are continuous and unimodal when parameters are varying in between interval endpoints. The theory applies to common important discrete random Variables including Binomial Variable, Poisson Variable, negative Binomial Variable and hypergeometrical random Variable. The theory can be used to make relevant statistical inference more rigorous and less conservative.

  • minimum coverage probability of confidence intervals
    2007
    Co-Authors: Xinjia Chen
    Abstract:

    In this paper, we develop a general theory on the coverage probability of random intervals defined in terms of discrete random Variables with continuous parameter spaces. The theory shows that the minimum coverage probabilities of random intervals with respect to corresponding parameters are achieved at discrete finite sets and that the coverage probabilities are continuous and unimodal when parameters are varying in between interval endpoints. The theory applies to common important discrete random Variables including Binomial Variable, Poisson Variable, negative Binomial Variable and hypergeometrical random Variable. The theory can be used to make relevant statistical inference more rigorous and less conservative.

Lawrence D. Brown - One of the best experts on this subject based on the ideXlab platform.

  • In-season prediction of batting averages: A field test of empirical Bayes and Bayes methodologies
    The Annals of Applied Statistics, 2008
    Co-Authors: Lawrence D. Brown
    Abstract:

    Batting average is one of the principle performance measures for an individual baseball player. It is natural to statistically model this as a Binomial-Variable proportion, with a given (observed) number of qualifying attempts (called ``at-bats''), an observed number of successes (``hits'') distributed according to the Binomial distribution, and with a true (but unknown) value of $p_i$ that represents the player's latent ability. This is a common data structure in many statistical applications; and so the methodological study here has implications for such a range of applications. We look at batting records for each Major League player over the course of a single season (2005). The primary focus is on using only the batting records from an earlier part of the season (e.g., the first 3 months) in order to estimate the batter's latent ability, $p_i$, and consequently, also to predict their batting-average performance for the remainder of the season. Since we are using a season that has already concluded, we can then validate our estimation performance by comparing the estimated values to the actual values for the remainder of the season. The prediction methods to be investigated are motivated from empirical Bayes and hierarchical Bayes interpretations. A newly proposed nonparametric empirical Bayes procedure performs particularly well in the basic analysis of the full data set, though less well with analyses involving more homogeneous subsets of the data. In those more homogeneous situations better performance is obtained from appropriate versions of more familiar methods. In all situations the poorest performing choice is the na\"{{\i}}ve predictor which directly uses the current average to predict the future average.

  • in season prediction of batting averages a field test of empirical bayes and bayes methodologies
    The Annals of Applied Statistics, 2008
    Co-Authors: Lawrence D. Brown
    Abstract:

    Batting average is one of the principle performance measures for an individual baseball player. It is natural to statistically model this as a Binomial-Variable proportion, with a given (observed) number of qualifying attempts (called “at-bats”), an observed number of successes (“hits”) distributed according to the Binomial distribution, and with a true (but unknown) value of pi that represents the player’s latent ability. This is a common data structure in many statistical applications; and so the methodological study here has implications for such a range of applications. We look at batting records for each Major League player over the course of a single season (2005). The primary focus is on using only the batting records from an earlier part of the season (e.g., the first 3 months) in order to estimate the batter’s latent ability, pi, and consequently, also to predict their batting-average performance for the remainder of the season. Since we are using a season that has already concluded, we can then validate our estimation performance by comparing the estimated values to the actual values for the remainder of the season. The prediction methods to be investigated are motivated from empirical Bayes and hierarchical Bayes interpretations. A newly proposed nonparametric empirical Bayes procedure performs particularly well in the basic analysis of the full data set, though less well with analyses involving more homogeneous subsets of the data. In those more homogeneous situations better performance is obtained from appropriate versions of more familiar methods. In all situations the poorest performing choice is the naive predictor which directly uses the current average to predict the future average. One feature of all the statistical methodologies here is the preliminary use of a new form of variance stabilizing transformation in order to transform the Binomial data problem into a somewhat more familiar structure involving (approximately) Normal random Variables with known variances. This transformation technique is also used in the construction of a new empirical validation test of the Binomial model assumption that is the conceptual basis for all our analyses.

Schirneck Martin - One of the best experts on this subject based on the ideXlab platform.

  • The Minimization of Random Hypergraphs
    2020
    Co-Authors: Bläsius Thomas, Friedrich Tobias, Schirneck Martin
    Abstract:

    We investigate the maximum-entropy model $\mathcal{B}_{n,m,p}$ for random $n$-vertex, $m$-edge multi-hypergraphs with expected edge size $pn$. We show that the expected size of the minimization $\min(\mathcal{B}_{n,m,p})$, i.e., the number of inclusion-wise minimal edges of $\mathcal{B}_{n,m,p}$, undergoes a phase transition with respect to $m$. If $m$ is at most $1/(1-p)^{(1-p)n}$, then $\mathrm{E}[|\min(\mathcal{B}_{n,m,p})|]$ is of order $\Theta(m)$, while for $m \ge 1/(1-p)^{(1-p+\varepsilon)n}$ for any $\varepsilon > 0$, it is $\Theta( 2^{(\mathrm{H}(\alpha) + (1-\alpha) \log_2 p) n}/ \sqrt{n})$. Here, $\mathrm{H}$ denotes the binary entropy function and $\alpha = - (\log_{1-p} m)/n$. The result implies that the maximum expected number of minimal edges over all $m$ is $\Theta((1+p)^n/\sqrt{n})$. Our structural findings have algorithmic implications for minimizing an input hypergraph, which has applications in the profiling of relational databases as well as for the Orthogonal Vectors problem studied in fine-grained complexity. We make several technical contributions that are of independent interest in probability. First, we improve the Chernoff--Hoeffding theorem on the tail of the Binomial distribution. In detail, we show that for a Binomial Variable $Y \sim \operatorname{Bin}(n,p)$ and any $0 < x < p$, it holds that $\mathrm{P}[Y \le xn] = \Theta( 2^{-\!\mathrm{D}(x \,{\|}\, p) n}/\sqrt{n})$, where $\mathrm{D}$ is the binary Kullback--Leibler divergence between Bernoulli distributions. We give explicit upper and lower bounds on the constants hidden in the big-O notation that hold for all $n$. Secondly, we establish the fact that the probability of a set of cardinality $i$ being minimal after $m$ i.i.d. maximum-entropy trials exhibits a sharp threshold behavior at $i^* = n + \log_{1-p} m$.Comment: 24 pages, 2 figure

  • The Minimization of Random Hypergraphs
    2020
    Co-Authors: Bläsius Thomas, Friedrich Tobias, Schirneck Martin
    Abstract:

    We investigate the maximum-entropy model $\mathcal{B}_{n,m,p}$ for random $n$-vertex, $m$-edge multi-hypergraphs with expected edge size $pn$. We show that the expected size of the minimization of $\mathcal{B}_{n,m,p}$, i.e., the number of its inclusion-wise minimal edges, undergoes a phase transition with respect to $m$. If $m$ is at most $1/(1-p)^{(1-p)n}$, then the minimization is of size $\Theta(m)$. Beyond that point, for $\alpha$ such that $m = 1/(1-p)^{\alpha n}$ and $\mathrm{H}$ being the entropy function, it is $\Theta(1) \cdot \min\!\left(1, \, \frac{1}{(\alpha\,{-}\,(1-p)) \sqrt{(1\,{-}\,\alpha) n}}\right) \cdot 2^{(\mathrm{H}(\alpha) + (1-\alpha) \log_2 p) n}.$ This implies that the maximum expected size over all $m$ is $\Theta((1+p)^n/\sqrt{n})$. Our structural findings have algorithmic implications for minimizing an input hypergraph, which in turn has applications in the profiling of relational databases as well as for the Orthogonal Vectors problem studied in fine-grained complexity. The main technical tool is an improvement of the Chernoff--Hoeffding inequality, which we make tight up to constant factors. We show that for a Binomial Variable $X \sim \mathrm{Bin}(n,p)$ and real number $0 < x \le p$, it holds that $\mathrm{P}[X \le xn] = \Theta(1) \cdot \min\!\left(1, \, \frac{1}{(p-x) \sqrt{xn}}\right) \cdot 2^{-\!\mathrm{D}(x \,{\|}\, p) n}$, where $\mathrm{D}$ denotes the Kullback--Leibler divergence between Bernoulli distributions. The result remains true if $x$ depends on $n$ as long as it is bounded away from $0$.Comment: 28 pages, 2 figures; Changes: Binomial characterization unified, improvement of the Chernoff-Hoeffding theorem extended to case x -->

  • The Minimization of Random Hypergraphs
    LIPIcs - Leibniz International Proceedings in Informatics. 28th Annual European Symposium on Algorithms (ESA 2020), 2020
    Co-Authors: Friedrich Tobias, Schirneck Martin
    Abstract:

    We investigate the maximum-entropy model B_{n,m,p} for random n-vertex, m-edge multi-hypergraphs with expected edge size pn. We show that the expected size of the minimization min(B_{n,m,p}), i.e., the number of inclusion-wise minimal edges of B_{n,m,p}, undergoes a phase transition with respect to m. If m is at most 1/(1-p)^{(1-p)n}, then E[|min(B_{n,m,p})|] is of order ?(m), while for m ? 1/(1-p)^{(1-p+?)n} for any ? > 0, it is ?(2^{(H(?) + (1-?) log? p) n}/?n). Here, H denotes the binary entropy function and ? = - (log_{1-p} m)/n. The result implies that the maximum expected number of minimal edges over all m is ?((1+p)?/?n). Our structural findings have algorithmic implications for minimizing an input hypergraph. This has applications in the profiling of relational databases as well as for the Orthogonal Vectors problem studied in fine-grained complexity. We make several technical contributions that are of independent interest in probability. First, we improve the Chernoff-Hoeffding theorem on the tail of the Binomial distribution. In detail, we show that for a Binomial Variable Y ? Bin(n,p) and any 0 < x < p, it holds that P[Y ? xn] = ?(2^{-D(x?p) n}/?n), where D is the binary Kullback-Leibler divergence between Bernoulli distributions. We give explicit upper and lower bounds on the constants hidden in the big-O notation that hold for all n. Secondly, we establish the fact that the probability of a set of cardinality i being minimal after m i.i.d. maximum-entropy trials exhibits a sharp threshold behavior at i^* = n + log_{1-p} m

Friedrich Tobias - One of the best experts on this subject based on the ideXlab platform.

  • The Minimization of Random Hypergraphs
    2020
    Co-Authors: Bläsius Thomas, Friedrich Tobias, Schirneck Martin
    Abstract:

    We investigate the maximum-entropy model $\mathcal{B}_{n,m,p}$ for random $n$-vertex, $m$-edge multi-hypergraphs with expected edge size $pn$. We show that the expected size of the minimization $\min(\mathcal{B}_{n,m,p})$, i.e., the number of inclusion-wise minimal edges of $\mathcal{B}_{n,m,p}$, undergoes a phase transition with respect to $m$. If $m$ is at most $1/(1-p)^{(1-p)n}$, then $\mathrm{E}[|\min(\mathcal{B}_{n,m,p})|]$ is of order $\Theta(m)$, while for $m \ge 1/(1-p)^{(1-p+\varepsilon)n}$ for any $\varepsilon > 0$, it is $\Theta( 2^{(\mathrm{H}(\alpha) + (1-\alpha) \log_2 p) n}/ \sqrt{n})$. Here, $\mathrm{H}$ denotes the binary entropy function and $\alpha = - (\log_{1-p} m)/n$. The result implies that the maximum expected number of minimal edges over all $m$ is $\Theta((1+p)^n/\sqrt{n})$. Our structural findings have algorithmic implications for minimizing an input hypergraph, which has applications in the profiling of relational databases as well as for the Orthogonal Vectors problem studied in fine-grained complexity. We make several technical contributions that are of independent interest in probability. First, we improve the Chernoff--Hoeffding theorem on the tail of the Binomial distribution. In detail, we show that for a Binomial Variable $Y \sim \operatorname{Bin}(n,p)$ and any $0 < x < p$, it holds that $\mathrm{P}[Y \le xn] = \Theta( 2^{-\!\mathrm{D}(x \,{\|}\, p) n}/\sqrt{n})$, where $\mathrm{D}$ is the binary Kullback--Leibler divergence between Bernoulli distributions. We give explicit upper and lower bounds on the constants hidden in the big-O notation that hold for all $n$. Secondly, we establish the fact that the probability of a set of cardinality $i$ being minimal after $m$ i.i.d. maximum-entropy trials exhibits a sharp threshold behavior at $i^* = n + \log_{1-p} m$.Comment: 24 pages, 2 figure

  • The Minimization of Random Hypergraphs
    2020
    Co-Authors: Bläsius Thomas, Friedrich Tobias, Schirneck Martin
    Abstract:

    We investigate the maximum-entropy model $\mathcal{B}_{n,m,p}$ for random $n$-vertex, $m$-edge multi-hypergraphs with expected edge size $pn$. We show that the expected size of the minimization of $\mathcal{B}_{n,m,p}$, i.e., the number of its inclusion-wise minimal edges, undergoes a phase transition with respect to $m$. If $m$ is at most $1/(1-p)^{(1-p)n}$, then the minimization is of size $\Theta(m)$. Beyond that point, for $\alpha$ such that $m = 1/(1-p)^{\alpha n}$ and $\mathrm{H}$ being the entropy function, it is $\Theta(1) \cdot \min\!\left(1, \, \frac{1}{(\alpha\,{-}\,(1-p)) \sqrt{(1\,{-}\,\alpha) n}}\right) \cdot 2^{(\mathrm{H}(\alpha) + (1-\alpha) \log_2 p) n}.$ This implies that the maximum expected size over all $m$ is $\Theta((1+p)^n/\sqrt{n})$. Our structural findings have algorithmic implications for minimizing an input hypergraph, which in turn has applications in the profiling of relational databases as well as for the Orthogonal Vectors problem studied in fine-grained complexity. The main technical tool is an improvement of the Chernoff--Hoeffding inequality, which we make tight up to constant factors. We show that for a Binomial Variable $X \sim \mathrm{Bin}(n,p)$ and real number $0 < x \le p$, it holds that $\mathrm{P}[X \le xn] = \Theta(1) \cdot \min\!\left(1, \, \frac{1}{(p-x) \sqrt{xn}}\right) \cdot 2^{-\!\mathrm{D}(x \,{\|}\, p) n}$, where $\mathrm{D}$ denotes the Kullback--Leibler divergence between Bernoulli distributions. The result remains true if $x$ depends on $n$ as long as it is bounded away from $0$.Comment: 28 pages, 2 figures; Changes: Binomial characterization unified, improvement of the Chernoff-Hoeffding theorem extended to case x -->

  • The Minimization of Random Hypergraphs
    LIPIcs - Leibniz International Proceedings in Informatics. 28th Annual European Symposium on Algorithms (ESA 2020), 2020
    Co-Authors: Friedrich Tobias, Schirneck Martin
    Abstract:

    We investigate the maximum-entropy model B_{n,m,p} for random n-vertex, m-edge multi-hypergraphs with expected edge size pn. We show that the expected size of the minimization min(B_{n,m,p}), i.e., the number of inclusion-wise minimal edges of B_{n,m,p}, undergoes a phase transition with respect to m. If m is at most 1/(1-p)^{(1-p)n}, then E[|min(B_{n,m,p})|] is of order ?(m), while for m ? 1/(1-p)^{(1-p+?)n} for any ? > 0, it is ?(2^{(H(?) + (1-?) log? p) n}/?n). Here, H denotes the binary entropy function and ? = - (log_{1-p} m)/n. The result implies that the maximum expected number of minimal edges over all m is ?((1+p)?/?n). Our structural findings have algorithmic implications for minimizing an input hypergraph. This has applications in the profiling of relational databases as well as for the Orthogonal Vectors problem studied in fine-grained complexity. We make several technical contributions that are of independent interest in probability. First, we improve the Chernoff-Hoeffding theorem on the tail of the Binomial distribution. In detail, we show that for a Binomial Variable Y ? Bin(n,p) and any 0 < x < p, it holds that P[Y ? xn] = ?(2^{-D(x?p) n}/?n), where D is the binary Kullback-Leibler divergence between Bernoulli distributions. We give explicit upper and lower bounds on the constants hidden in the big-O notation that hold for all n. Secondly, we establish the fact that the probability of a set of cardinality i being minimal after m i.i.d. maximum-entropy trials exhibits a sharp threshold behavior at i^* = n + log_{1-p} m