The Experts below are selected from a list of 22344 Experts worldwide ranked by ideXlab platform
Jose C Principe - One of the best experts on this subject based on the ideXlab platform.
-
deep deterministic information bottleneck with matrix based Entropy Functional
International Conference on Acoustics Speech and Signal Processing, 2021Co-Authors: Jose C PrincipeAbstract:We introduce the matrix-based Renyi’s α-order Entropy Functional to parameterize Tishby et al. information bottleneck (IB) principle [1] with a neural network. We term our methodology Deep Deterministic Information Bottleneck (DIB), as it avoids variational inference and distribution assumption. We show that deep neural networks trained with DIB outperform the variational objective counterpart and those that are trained with other forms of regularization, in terms of generalization performance and robustness to adversarial attack. Code available at https://github.com/yuxi120407/DIB.
-
interpretable fault detection using projections of mutual information matrix
Journal of The Franklin Institute-engineering and Applied Mathematics, 2021Co-Authors: Chenglin Wen, Jose C PrincipeAbstract:Abstract This paper presents a novel mutual information (MI) matrix based method for fault detection. Given a m -dimensional fault process, the MI matrix is a m × m matrix in which the ( i , j ) -th entry measures the MI values between the i -th dimension and the j -th dimension variables. We introduce the recently proposed matrix-based Renyi’s α -Entropy Functional to estimate MI values in each entry of the MI matrix. The new estimator avoids density estimation and it operates on the eigenspectrum of a (normalized) symmetric positive definite (SPD) matrix, which makes it well suited for industrial process. We combine different orders of statistics of the transformed components (TCs) extracted from the MI matrix to constitute the detection index, and derive a simple similarity index to monitor the changes of characteristics of the underlying process in consecutive windows. We term the overall methodology “projections of mutual information matrix” (PMIM). Experiments on both synthetic data and the benchmark Tennessee Eastman process demonstrate the interpretability of PMIM in identifying the root variables that cause the faults, and its superiority in detecting the occurrence of faults in terms of the improved fault detection rate (FDR) and the lowest false alarm rate (FAR). The advantages of PMIM is also less sensitive to hyper-parameters. The advantages of PMIM is also less sensitive to hyper-parameters. Code of PMIM is available at https://github.com/SJYuCNEL/Fault_detection_PMIM .
-
measuring dependence with matrix based Entropy Functional
National Conference on Artificial Intelligence, 2021Co-Authors: Francesco Alesiani, Robert Jenssen, Jose C PrincipeAbstract:Measuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic dependence measures into a higher-level perspective by the Shearer's inequality. Based on our generalization, we then propose two measures, namely the matrix-based normalized total correlation ($T_\alpha^*$) and the matrix-based normalized dual total correlation ($D_\alpha^*$), to quantify the dependence of multiple variables in arbitrary dimensional space, without explicit estimation of the underlying data distributions. We show that our measures are differentiable and statistically more powerful than prevalent ones. We also show the impact of our measures in four different machine learning problems, namely the gene regulatory network inference, the robust machine learning under covariate shift and non-Gaussian noises, the subspace outlier detection, and the understanding of the learning dynamics of convolutional neural networks (CNNs), to demonstrate their utilities, advantages, as well as implications to those problems. Code of our dependence measure is available at: this https URL
-
multivariate extension of matrix based renyi s alpha α order Entropy Functional
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020Co-Authors: Luis Gonzalo Sanchez Giraldo, Robert Jenssen, Jose C PrincipeAbstract:The matrix-based Renyi's $\alpha$ α -order Entropy Functional was recently introduced using the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the current theory in the matrix-based Renyi's $\alpha$ α -order Entropy Functional only defines the Entropy of a single variable or mutual information between two random variables. In information theory and machine learning communities, one is also frequently interested in multivariate information quantities, such as the multivariate joint Entropy and different interactive quantities among multiple variables. In this paper, we first define the matrix-based Renyi's $\alpha$ α -order joint Entropy among multiple variables. We then show how this definition can ease the estimation of various information quantities that measure the interactions among multiple variables, such as interactive information and total correlation. We finally present an application to feature selection to show how our definition provides a simple yet powerful way to estimate a widely-acknowledged intractable quantity from data. A real example on hyperspectral image (HSI) band selection is also provided.
-
interpretable fault detection using projections of mutual information matrix
arXiv: Signal Processing, 2020Co-Authors: Chenglin Wen, Jose C PrincipeAbstract:This paper presents a novel mutual information (MI) matrix based method for fault detection. Given a $m$-dimensional fault process, the MI matrix is a $m \times m$ matrix in which the $(i,j)$-th entry measures the MI values between the $i$-th dimension and the $j$-th dimension variables. We introduce the recently proposed matrix-based R\'enyi's $\alpha$-Entropy Functional to estimate MI values in each entry of the MI matrix. The new estimator avoids density estimation and it operates on the eigenspectrum of a (normalized) symmetric positive definite (SPD) matrix, which makes it well suited for industrial process. We combine different orders of statistics of the transformed components (TCs) extracted from the MI matrix to constitute the detection index, and derive a simple similarity index to monitor the changes of characteristics of the underlying process in consecutive windows. We term the overall methodology "projections of mutual information matrix" (PMIM). Experiments on both synthetic data and the benchmark Tennessee Eastman process demonstrate the interpretability of PMIM in identifying the root variables that cause the faults, and its superiority in detecting the occurrence of faults in terms of the improved fault detection rate (FDR) and the lowest false alarm rate (FAR). The advantages of PMIM is also less sensitive to hyper-parameters. The advantages of PMIM is also less sensitive to hyper-parameters. Code of PMIM is available at https://github.com/SJYuCNEL/Fault_detection_PMIM
Tsachy Weissman - One of the best experts on this subject based on the ideXlab platform.
-
maximum likelihood estimation of Functionals of discrete distributions
IEEE Transactions on Information Theory, 2017Co-Authors: Jiantao Jiao, Yanjun Han, Kartik Venkat, Tsachy WeissmanAbstract:We consider the problem of estimating Functionals of discrete distributions, and focus on a tight (up to universal multiplicative constants for each specific Functional) nonasymptotic analysis of the worst case squared error risk of widely used estimators. We apply concentration inequalities to analyze the random fluctuation of these estimators around their expectations and the theory of approximation using positive linear operators to analyze the deviation of their expectations from the true Functional, namely their bias . We explicitly characterize the worst case squared error risk incurred by the maximum likelihood estimator (MLE) in estimating the Shannon Entropy $H(P) = \sum _{i = 1}^{S} -p_{i} \ln p_{i}$ , and the power sum $F_\alpha (P) = \sum _{i = 1}^{S} p_{i}^\alpha ,\alpha >0$ , up to universal multiplicative constants for each fixed Functional, for any alphabet size $S\leq \infty $ and sample size $n$ for which the risk may vanish. As a corollary, for Shannon Entropy estimation, we show that it is necessary and sufficient to have $n \gg S$ observations for the MLE to be consistent. In addition, we establish that it is necessary and sufficient to consider $n \gg S^{1/\alpha }$ samples for the MLE to consistently estimate $F_\alpha (P), 0 . The minimax rate-optimal estimators for both problems require $S/\ln S$ and $S^{1/\alpha }/\ln S$ samples, which implies that the MLE has a strictly sub-optimal sample complexity. When $1 , we show that the worst case squared error rate of convergence for the MLE is $n^{-2(\alpha -1)}$ for infinite alphabet size, while the minimax squared error rate is $(n\ln n)^{-2(\alpha -1)}$ . When $\alpha \geq 3/2$ , the MLE achieves the minimax optimal rate $n^{-1}$ regardless of the alphabet size. As an application of the general theory, we analyze the Dirichlet prior smoothing techniques for Shannon Entropy estimation. In this context, one approach is to plug-in the Dirichlet prior smoothed distribution into the Entropy Functional, while the other one is to calculate the Bayes estimator for Entropy under the Dirichlet prior for squared error, which is the conditional expectation. We show that in general such estimators do not improve over the maximum likelihood estimator. No matter how we tune the parameters in the Dirichlet prior, this approach cannot achieve the minimax rates in Entropy estimation. The performance of the minimax rate-optimal estimator with $n$ samples is essentially at least as good as that of Dirichlet smoothed Entropy estimators with $n\ln n$ samples.
-
does dirichlet prior smoothing solve the shannon Entropy estimation problem
International Symposium on Information Theory, 2015Co-Authors: Yanjun Han, Jiantao Jiao, Tsachy WeissmanAbstract:The Dirichlet prior is widely used in estimating discrete distributions and Functionals of discrete distributions. In terms of Shannon Entropy estimation, one approach is to plug-in the Dirichlet prior smoothed distribution into the Entropy Functional, while the other one is to calculate the Bayes estimator for Entropy under the Dirichlet prior for squared error, which is the conditional expectation. We show that in general they do not improve over the maximum likelihood estimator, which plugs-in the empirical distribution into the Entropy Functional. No matter how we tune the parameters in the Dirichlet prior, this approach cannot achieve the minimax rates in Entropy estimation, as recently characterized by Jiao, Venkat, Han, and Weissman [1], and Wu and Yang [2]. The performance of the minimax rate-optimal estimator with n samples is essentially at least as good as that of the Dirichlet smoothed Entropy estimators with n ln n samples. We harness the theory of approximation using positive linear operators for analyzing the bias of plug-in estimators for general Functionals under arbitrary statistical models, thereby further consolidating the interplay between these two fields, which was thoroughly exploited by Jiao, Venkat, Han, and Weissman [3] in estimating various Functionals of discrete distributions. We establish new results in approximation theory, and apply them to analyze the bias of the Dirichlet prior smoothed plug-in Entropy estimator. This interplay between bias analysis and approximation theory is of relevance and consequence far beyond the specific problem setting in this paper.
-
does dirichlet prior smoothing solve the shannon Entropy estimation problem
arXiv: Information Theory, 2015Co-Authors: Yanjun Han, Jiantao Jiao, Tsachy WeissmanAbstract:The Dirichlet prior is widely used in estimating discrete distributions and Functionals of discrete distributions. In terms of Shannon Entropy estimation, one approach is to plug-in the Dirichlet prior smoothed distribution into the Entropy Functional, while the other one is to calculate the Bayes estimator for Entropy under the Dirichlet prior for squared error, which is the conditional expectation. We show that in general they do \emph{not} improve over the maximum likelihood estimator, which plugs-in the empirical distribution into the Entropy Functional. No matter how we tune the parameters in the Dirichlet prior, this approach cannot achieve the minimax rates in Entropy estimation, as recently characterized by Jiao, Venkat, Han, and Weissman, and Wu and Yang. The performance of the minimax rate-optimal estimator with $n$ samples is essentially \emph{at least} as good as that of the Dirichlet smoothed Entropy estimators with $n\ln n$ samples. We harness the theory of approximation using positive linear operators for analyzing the bias of plug-in estimators for general Functionals under arbitrary statistical models, thereby further consolidating the interplay between these two fields, which was thoroughly developed and exploited by Jiao, Venkat, Han, and Weissman. We establish new results in approximation theory, and apply them to analyze the bias of the Dirichlet prior smoothed plug-in Entropy estimator. This interplay between bias analysis and approximation theory is of relevance and consequence far beyond the specific problem setting in this paper.
Luiz G Ferreira - One of the best experts on this subject based on the ideXlab platform.
-
evaluating and improving the cluster variation method Entropy Functional for ising alloys
Journal of Chemical Physics, 1998Co-Authors: Luiz G Ferreira, Christopher M Wolverton, Alex ZungerAbstract:The success of the “cluster variation method” (CVM) in reproducing quite accurately the free energies of Monte Carlo (MC) calculations on Ising models is explained in terms of identifying a cancellation of errors: We show that the CVM produces correlation functions that are too close to zero, which leads to an overestimation of the exact energy, E, and at the same time, to an underestimation of −TS, so the free energy F=E−TS is more accurate than either of its parts. This insight explains a problem with “hybrid methods” using MC correlation functions in the CVM Entropy expression: They give exact energies E and do not give significantly improved −TS relative to CVM, so they do not benefit from the above noted cancellation of errors. Additionally, hybrid methods suffer from the difficulty of adequately accounting for both ordered and disordered phases in a consistent way. A different technique, the “entropic Monte Carlo” (EMC), is shown here to provide a means for critically evaluating the CVM Entropy. Ins...
-
evaluating and improving the cluster variation method Entropy Functional for ising alloys
arXiv: Materials Science, 1998Co-Authors: Luiz G Ferreira, Chris Wolverton, Alex ZungerAbstract:The success of the "Cluster Variation Method" (CVM) in reproducing quite accurately the free energies of Monte Carlo (MC) calculations on Ising models is explained in terms of identifying a cancellation of errors: We show that the CVM produces correlation functions that are too close to zero, which leads to an overestimation of the exact energy, E, and at the same time, to an underestimation of -TS, so the free energy F=E-TS is more accurate than either of its parts. This insight explains a problem with "hybrid methods" using MC correlation functions in the CVM Entropy expression: They give exact energies E and do not give significantly improved -TS relative to CVM, so they do not benefit from the above noted cancellation of errors. Additionally, "hybrid methods" suffer from the difficulty of adequately accounting for both ordered and disordered phases in a consistent way. A different technique, the "Entropic Monte Carlo" (EMC), is shown here to provide a means for critically evaluating the CVM Entropy. Inspired by EMC results, we find a universal and simple correction to the CVM Entropy which produces individual components of the free energy with MC accuracy, but is computationally much less expensive than either MC thermodynamic integration or EMC.
Yanjun Han - One of the best experts on this subject based on the ideXlab platform.
-
maximum likelihood estimation of Functionals of discrete distributions
IEEE Transactions on Information Theory, 2017Co-Authors: Jiantao Jiao, Yanjun Han, Kartik Venkat, Tsachy WeissmanAbstract:We consider the problem of estimating Functionals of discrete distributions, and focus on a tight (up to universal multiplicative constants for each specific Functional) nonasymptotic analysis of the worst case squared error risk of widely used estimators. We apply concentration inequalities to analyze the random fluctuation of these estimators around their expectations and the theory of approximation using positive linear operators to analyze the deviation of their expectations from the true Functional, namely their bias . We explicitly characterize the worst case squared error risk incurred by the maximum likelihood estimator (MLE) in estimating the Shannon Entropy $H(P) = \sum _{i = 1}^{S} -p_{i} \ln p_{i}$ , and the power sum $F_\alpha (P) = \sum _{i = 1}^{S} p_{i}^\alpha ,\alpha >0$ , up to universal multiplicative constants for each fixed Functional, for any alphabet size $S\leq \infty $ and sample size $n$ for which the risk may vanish. As a corollary, for Shannon Entropy estimation, we show that it is necessary and sufficient to have $n \gg S$ observations for the MLE to be consistent. In addition, we establish that it is necessary and sufficient to consider $n \gg S^{1/\alpha }$ samples for the MLE to consistently estimate $F_\alpha (P), 0 . The minimax rate-optimal estimators for both problems require $S/\ln S$ and $S^{1/\alpha }/\ln S$ samples, which implies that the MLE has a strictly sub-optimal sample complexity. When $1 , we show that the worst case squared error rate of convergence for the MLE is $n^{-2(\alpha -1)}$ for infinite alphabet size, while the minimax squared error rate is $(n\ln n)^{-2(\alpha -1)}$ . When $\alpha \geq 3/2$ , the MLE achieves the minimax optimal rate $n^{-1}$ regardless of the alphabet size. As an application of the general theory, we analyze the Dirichlet prior smoothing techniques for Shannon Entropy estimation. In this context, one approach is to plug-in the Dirichlet prior smoothed distribution into the Entropy Functional, while the other one is to calculate the Bayes estimator for Entropy under the Dirichlet prior for squared error, which is the conditional expectation. We show that in general such estimators do not improve over the maximum likelihood estimator. No matter how we tune the parameters in the Dirichlet prior, this approach cannot achieve the minimax rates in Entropy estimation. The performance of the minimax rate-optimal estimator with $n$ samples is essentially at least as good as that of Dirichlet smoothed Entropy estimators with $n\ln n$ samples.
-
does dirichlet prior smoothing solve the shannon Entropy estimation problem
International Symposium on Information Theory, 2015Co-Authors: Yanjun Han, Jiantao Jiao, Tsachy WeissmanAbstract:The Dirichlet prior is widely used in estimating discrete distributions and Functionals of discrete distributions. In terms of Shannon Entropy estimation, one approach is to plug-in the Dirichlet prior smoothed distribution into the Entropy Functional, while the other one is to calculate the Bayes estimator for Entropy under the Dirichlet prior for squared error, which is the conditional expectation. We show that in general they do not improve over the maximum likelihood estimator, which plugs-in the empirical distribution into the Entropy Functional. No matter how we tune the parameters in the Dirichlet prior, this approach cannot achieve the minimax rates in Entropy estimation, as recently characterized by Jiao, Venkat, Han, and Weissman [1], and Wu and Yang [2]. The performance of the minimax rate-optimal estimator with n samples is essentially at least as good as that of the Dirichlet smoothed Entropy estimators with n ln n samples. We harness the theory of approximation using positive linear operators for analyzing the bias of plug-in estimators for general Functionals under arbitrary statistical models, thereby further consolidating the interplay between these two fields, which was thoroughly exploited by Jiao, Venkat, Han, and Weissman [3] in estimating various Functionals of discrete distributions. We establish new results in approximation theory, and apply them to analyze the bias of the Dirichlet prior smoothed plug-in Entropy estimator. This interplay between bias analysis and approximation theory is of relevance and consequence far beyond the specific problem setting in this paper.
-
does dirichlet prior smoothing solve the shannon Entropy estimation problem
arXiv: Information Theory, 2015Co-Authors: Yanjun Han, Jiantao Jiao, Tsachy WeissmanAbstract:The Dirichlet prior is widely used in estimating discrete distributions and Functionals of discrete distributions. In terms of Shannon Entropy estimation, one approach is to plug-in the Dirichlet prior smoothed distribution into the Entropy Functional, while the other one is to calculate the Bayes estimator for Entropy under the Dirichlet prior for squared error, which is the conditional expectation. We show that in general they do \emph{not} improve over the maximum likelihood estimator, which plugs-in the empirical distribution into the Entropy Functional. No matter how we tune the parameters in the Dirichlet prior, this approach cannot achieve the minimax rates in Entropy estimation, as recently characterized by Jiao, Venkat, Han, and Weissman, and Wu and Yang. The performance of the minimax rate-optimal estimator with $n$ samples is essentially \emph{at least} as good as that of the Dirichlet smoothed Entropy estimators with $n\ln n$ samples. We harness the theory of approximation using positive linear operators for analyzing the bias of plug-in estimators for general Functionals under arbitrary statistical models, thereby further consolidating the interplay between these two fields, which was thoroughly developed and exploited by Jiao, Venkat, Han, and Weissman. We establish new results in approximation theory, and apply them to analyze the bias of the Dirichlet prior smoothed plug-in Entropy estimator. This interplay between bias analysis and approximation theory is of relevance and consequence far beyond the specific problem setting in this paper.
Alex Zunger - One of the best experts on this subject based on the ideXlab platform.
-
evaluating and improving the cluster variation method Entropy Functional for ising alloys
Journal of Chemical Physics, 1998Co-Authors: Luiz G Ferreira, Christopher M Wolverton, Alex ZungerAbstract:The success of the “cluster variation method” (CVM) in reproducing quite accurately the free energies of Monte Carlo (MC) calculations on Ising models is explained in terms of identifying a cancellation of errors: We show that the CVM produces correlation functions that are too close to zero, which leads to an overestimation of the exact energy, E, and at the same time, to an underestimation of −TS, so the free energy F=E−TS is more accurate than either of its parts. This insight explains a problem with “hybrid methods” using MC correlation functions in the CVM Entropy expression: They give exact energies E and do not give significantly improved −TS relative to CVM, so they do not benefit from the above noted cancellation of errors. Additionally, hybrid methods suffer from the difficulty of adequately accounting for both ordered and disordered phases in a consistent way. A different technique, the “entropic Monte Carlo” (EMC), is shown here to provide a means for critically evaluating the CVM Entropy. Ins...
-
evaluating and improving the cluster variation method Entropy Functional for ising alloys
arXiv: Materials Science, 1998Co-Authors: Luiz G Ferreira, Chris Wolverton, Alex ZungerAbstract:The success of the "Cluster Variation Method" (CVM) in reproducing quite accurately the free energies of Monte Carlo (MC) calculations on Ising models is explained in terms of identifying a cancellation of errors: We show that the CVM produces correlation functions that are too close to zero, which leads to an overestimation of the exact energy, E, and at the same time, to an underestimation of -TS, so the free energy F=E-TS is more accurate than either of its parts. This insight explains a problem with "hybrid methods" using MC correlation functions in the CVM Entropy expression: They give exact energies E and do not give significantly improved -TS relative to CVM, so they do not benefit from the above noted cancellation of errors. Additionally, "hybrid methods" suffer from the difficulty of adequately accounting for both ordered and disordered phases in a consistent way. A different technique, the "Entropic Monte Carlo" (EMC), is shown here to provide a means for critically evaluating the CVM Entropy. Inspired by EMC results, we find a universal and simple correction to the CVM Entropy which produces individual components of the free energy with MC accuracy, but is computationally much less expensive than either MC thermodynamic integration or EMC.