The Experts below are selected from a list of 88056 Experts worldwide ranked by ideXlab platform
Srinivas Reddy Geedipally - One of the best experts on this subject based on the ideXlab platform.
-
a semiparametric negative binomial Generalized Linear Model for Modeling over dispersed count data with a heavy tail characteristics and applications to crash data
Accident Analysis & Prevention, 2016Co-Authors: Mohammadali Shirazi, Soma S Dhavala, Dominique Lord, Srinivas Reddy GeedipallyAbstract:Crash data can often be characterized by over-dispersion, heavy (long) tail and many observations with the value zero. Over the last few years, a small number of researchers have started developing and applying novel and innovative multi-parameter Models to analyze such data. These multi-parameter Models have been proposed for overcoming the limitations of the traditional negative binomial (NB) Model, which cannot handle this kind of data efficiently. The research documented in this paper continues the work related to multi-parameter Models. The objective of this paper is to document the development and application of a flexible NB Generalized Linear Model with randomly distributed mixed effects characterized by the Dirichlet process (NB-DP) to Model crash data. The objective of the study was accomplished using two datasets. The new Model was compared to the NB and the recently introduced Model based on the mixture of the NB and Lindley (NB-L) distributions. Overall, the research study shows that the NB-DP Model offers a better performance than the NB Model once data are over-dispersed and have a heavy tail. The NB-DP performed better than the NB-L when the dataset has a heavy tail, but a smaller percentage of zeros. However, both Models performed similarly when the dataset contained a large amount of zeros. In addition to a greater flexibility, the NB-DP provides a clustering by-product that allows the safety analyst to better understand the characteristics of the data, such as the identification of outliers and sources of dispersion.
-
the poisson weibull Generalized Linear Model for analyzing motor vehicle crash data
Safety Science, 2013Co-Authors: Lingzi Cheng, Srinivas Reddy Geedipally, Dominique LordAbstract:Abstract Over the last 20–30 years, there has been a significant amount of tools and statistical methods that have been proposed for analyzing crash data. Yet, the Poisson–gamma (PG) is still the most commonly used and widely acceptable Model. This paper documents the application of the Poisson–Weibull (PW) Generalized Linear Model (GLM) for Modeling motor vehicle crashes. The objectives of this study were to evaluate the application of the PW GLM for analyzing this kind of dataset and compare the results with the traditional PG Model. To accomplish the objectives of the study, the Modeling performance of the PW Model was first examined using a simulated dataset and then several PW and PG GLMs were developed and compared using two observed crash datasets. The results of this study show that the PW GLM performs as well as the PG GLM in terms of goodness-of-fit statistics.
-
the negative binomial lindley Generalized Linear Model characteristics and application using crash data
Accident Analysis & Prevention, 2012Co-Authors: Srinivas Reddy Geedipally, Dominique Lord, Soma S DhavalaAbstract:There has been a considerable amount of work devoted by transportation safety analysts to the development and application of new and innovative Models for analyzing crash data. One important characteristic about crash data that has been documented in the literature is related to datasets that contained a large amount of zeros and a long or heavy tail (which creates highly dispersed data). For such datasets, the number of sites where no crash is observed is so large that traditional distributions and regression Models, such as the Poisson and Poisson-gamma or negative binomial (NB) Models cannot be used efficiently. To overcome this problem, the NB-Lindley (NB-L) distribution has recently been introduced for analyzing count data that are characterized by excess zeros. The objective of this paper is to document the application of a NB Generalized Linear Model with Lindley mixed effects (NB-L GLM) for analyzing traffic crash data. The study objective was accomplished using simulated and observed datasets. The simulated dataset was used to show the general performance of the Model. The Model was then applied to two datasets based on observed data. One of the dataset was characterized by a large amount of zeros. The NB-L GLM was compared with the NB and zero-inflated Models. Overall, the research study shows that the NB-L GLM not only offers superior performance over the NB and zero-inflated Models when datasets are characterized by a large number of zeros and a long tail, but also when the crash dataset is highly dispersed.
-
application of the conway maxwell poisson Generalized Linear Model for analyzing motor vehicle crashes
Accident Analysis & Prevention, 2008Co-Authors: Dominique Lord, Seth D Guikema, Srinivas Reddy GeedipallyAbstract:Abstract This paper documents the application of the Conway–Maxwell–Poisson (COM-Poisson) Generalized Linear Model (GLM) for Modeling motor vehicle crashes. The COM-Poisson distribution, originally developed in 1962, has recently been re-introduced by statisticians for analyzing count data subjected to over- and under-dispersion. This innovative distribution is an extension of the Poisson distribution. The objectives of this study were to evaluate the application of the COM-Poisson GLM for analyzing motor vehicle crashes and compare the results with the traditional negative binomial (NB) Model. The comparison analysis was carried out using the most common functional forms employed by transportation safety analysts, which link crashes to the entering flows at intersections or on segments. To accomplish the objectives of the study, several NB and COM-Poisson GLMs were developed and compared using two datasets. The first dataset contained crash data collected at signalized four-legged intersections in Toronto, Ont. The second dataset included data collected for rural four-lane divided and undivided highways in Texas. Several methods were used to assess the statistical fit and predictive performance of the Models. The results of this study show that COM-Poisson GLMs perform as well as NB Models in terms of GOF statistics and predictive performance. Given the fact the COM-Poisson distribution can also handle under-dispersed data (while the NB distribution cannot or has difficulties converging), which have sometimes been observed in crash databases, the COM-Poisson GLM offers a better alternative over the NB Model for Modeling motor vehicle crashes, especially given the important limitations recently documented in the safety literature about the latter type of Model.
Dominique Lord - One of the best experts on this subject based on the ideXlab platform.
-
a semiparametric negative binomial Generalized Linear Model for Modeling over dispersed count data with a heavy tail characteristics and applications to crash data
Accident Analysis & Prevention, 2016Co-Authors: Mohammadali Shirazi, Soma S Dhavala, Dominique Lord, Srinivas Reddy GeedipallyAbstract:Crash data can often be characterized by over-dispersion, heavy (long) tail and many observations with the value zero. Over the last few years, a small number of researchers have started developing and applying novel and innovative multi-parameter Models to analyze such data. These multi-parameter Models have been proposed for overcoming the limitations of the traditional negative binomial (NB) Model, which cannot handle this kind of data efficiently. The research documented in this paper continues the work related to multi-parameter Models. The objective of this paper is to document the development and application of a flexible NB Generalized Linear Model with randomly distributed mixed effects characterized by the Dirichlet process (NB-DP) to Model crash data. The objective of the study was accomplished using two datasets. The new Model was compared to the NB and the recently introduced Model based on the mixture of the NB and Lindley (NB-L) distributions. Overall, the research study shows that the NB-DP Model offers a better performance than the NB Model once data are over-dispersed and have a heavy tail. The NB-DP performed better than the NB-L when the dataset has a heavy tail, but a smaller percentage of zeros. However, both Models performed similarly when the dataset contained a large amount of zeros. In addition to a greater flexibility, the NB-DP provides a clustering by-product that allows the safety analyst to better understand the characteristics of the data, such as the identification of outliers and sources of dispersion.
-
the poisson weibull Generalized Linear Model for analyzing motor vehicle crash data
Safety Science, 2013Co-Authors: Lingzi Cheng, Srinivas Reddy Geedipally, Dominique LordAbstract:Abstract Over the last 20–30 years, there has been a significant amount of tools and statistical methods that have been proposed for analyzing crash data. Yet, the Poisson–gamma (PG) is still the most commonly used and widely acceptable Model. This paper documents the application of the Poisson–Weibull (PW) Generalized Linear Model (GLM) for Modeling motor vehicle crashes. The objectives of this study were to evaluate the application of the PW GLM for analyzing this kind of dataset and compare the results with the traditional PG Model. To accomplish the objectives of the study, the Modeling performance of the PW Model was first examined using a simulated dataset and then several PW and PG GLMs were developed and compared using two observed crash datasets. The results of this study show that the PW GLM performs as well as the PG GLM in terms of goodness-of-fit statistics.
-
the negative binomial lindley Generalized Linear Model characteristics and application using crash data
Accident Analysis & Prevention, 2012Co-Authors: Srinivas Reddy Geedipally, Dominique Lord, Soma S DhavalaAbstract:There has been a considerable amount of work devoted by transportation safety analysts to the development and application of new and innovative Models for analyzing crash data. One important characteristic about crash data that has been documented in the literature is related to datasets that contained a large amount of zeros and a long or heavy tail (which creates highly dispersed data). For such datasets, the number of sites where no crash is observed is so large that traditional distributions and regression Models, such as the Poisson and Poisson-gamma or negative binomial (NB) Models cannot be used efficiently. To overcome this problem, the NB-Lindley (NB-L) distribution has recently been introduced for analyzing count data that are characterized by excess zeros. The objective of this paper is to document the application of a NB Generalized Linear Model with Lindley mixed effects (NB-L GLM) for analyzing traffic crash data. The study objective was accomplished using simulated and observed datasets. The simulated dataset was used to show the general performance of the Model. The Model was then applied to two datasets based on observed data. One of the dataset was characterized by a large amount of zeros. The NB-L GLM was compared with the NB and zero-inflated Models. Overall, the research study shows that the NB-L GLM not only offers superior performance over the NB and zero-inflated Models when datasets are characterized by a large number of zeros and a long tail, but also when the crash dataset is highly dispersed.
-
application of the conway maxwell poisson Generalized Linear Model for analyzing motor vehicle crashes
Accident Analysis & Prevention, 2008Co-Authors: Dominique Lord, Seth D Guikema, Srinivas Reddy GeedipallyAbstract:Abstract This paper documents the application of the Conway–Maxwell–Poisson (COM-Poisson) Generalized Linear Model (GLM) for Modeling motor vehicle crashes. The COM-Poisson distribution, originally developed in 1962, has recently been re-introduced by statisticians for analyzing count data subjected to over- and under-dispersion. This innovative distribution is an extension of the Poisson distribution. The objectives of this study were to evaluate the application of the COM-Poisson GLM for analyzing motor vehicle crashes and compare the results with the traditional negative binomial (NB) Model. The comparison analysis was carried out using the most common functional forms employed by transportation safety analysts, which link crashes to the entering flows at intersections or on segments. To accomplish the objectives of the study, several NB and COM-Poisson GLMs were developed and compared using two datasets. The first dataset contained crash data collected at signalized four-legged intersections in Toronto, Ont. The second dataset included data collected for rural four-lane divided and undivided highways in Texas. Several methods were used to assess the statistical fit and predictive performance of the Models. The results of this study show that COM-Poisson GLMs perform as well as NB Models in terms of GOF statistics and predictive performance. Given the fact the COM-Poisson distribution can also handle under-dispersed data (while the NB distribution cannot or has difficulties converging), which have sometimes been observed in crash databases, the COM-Poisson GLM offers a better alternative over the NB Model for Modeling motor vehicle crashes, especially given the important limitations recently documented in the safety literature about the latter type of Model.
Nicky J Welton - One of the best experts on this subject based on the ideXlab platform.
-
evidence synthesis for decision making 2 a Generalized Linear Modeling framework for pairwise and network meta analysis of randomized controlled trials
Medical Decision Making, 2013Co-Authors: Sofia Dias, Alex J Sutton, A E Ades, Nicky J WeltonAbstract:We set out a Generalized Linear Model framework for the synthesis of data from randomized controlled trials. A common Model is described, taking the form of a Linear regression for both fixed and random effects synthesis, which can be implemented with normal, binomial, Poisson, and multinomial data. The familiar logistic Model for meta-analysis with binomial data is a Generalized Linear Model with a logit link function, which is appropriate for probability outcomes. The same Linear regression framework can be applied to continuous outcomes, rate Models, competing risks, or ordered category outcomes by using other link functions, such as identity, log, complementary log-log, and probit link functions. The common core Model for the Linear predictor can be applied to pairwise meta-analysis, indirect comparisons, synthesis of multiarm trials, and mixed treatment comparisons, also known as network meta-analysis, without distinction. We take a Bayesian approach to estimation and provide WinBUGS program code for a Bayesian analysis using Markov chain Monte Carlo simulation. An advantage of this approach is that it is straightforward to extend to shared parameter Models where different randomized controlled trials report outcomes in different formats but from a common underlying Model. Use of the Generalized Linear Model framework allows us to present a unified account of how Models can be compared using the deviance information criterion and how goodness of fit can be assessed using the residual deviance. The approach is illustrated through a range of worked examples for commonly encountered evidence formats. Key words: Generalized Linear Model; network meta-analysis; indirect evidence; meta-analysis. (Med Decis Making 2013;33:607–617)
-
evidence synthesis for decision making 2 a Generalized Linear Modeling framework for pairwise and network meta analysis of randomized controlled trials
Medical Decision Making, 2013Co-Authors: Sofia Dias, Alex J Sutton, A E Ades, Nicky J WeltonAbstract:We set out a Generalized Linear Model framework for the synthesis of data from randomized controlled trials. A common Model is described, taking the form of a Linear regression for both fixed and random effects synthesis, which can be implemented with normal, binomial, Poisson, and multinomial data. The familiar logistic Model for meta-analysis with binomial data is a Generalized Linear Model with a logit link function, which is appropriate for probability outcomes. The same Linear regression framework can be applied to continuous outcomes, rate Models, competing risks, or ordered category outcomes by using other link functions, such as identity, log, complementary log-log, and probit link functions. The common core Model for the Linear predictor can be applied to pairwise meta-analysis, indirect comparisons, synthesis of multiarm trials, and mixed treatment comparisons, also known as network meta-analysis, without distinction. We take a Bayesian approach to estimation and provide WinBUGS program code for a Bayesian analysis using Markov chain Monte Carlo simulation. An advantage of this approach is that it is straightforward to extend to shared parameter Models where different randomized controlled trials report outcomes in different formats but from a common underlying Model. Use of the Generalized Linear Model framework allows us to present a unified account of how Models can be compared using the deviance information criterion and how goodness of fit can be assessed using the residual deviance. The approach is illustrated through a range of worked examples for commonly encountered evidence formats.
Soma S Dhavala - One of the best experts on this subject based on the ideXlab platform.
-
a semiparametric negative binomial Generalized Linear Model for Modeling over dispersed count data with a heavy tail characteristics and applications to crash data
Accident Analysis & Prevention, 2016Co-Authors: Mohammadali Shirazi, Soma S Dhavala, Dominique Lord, Srinivas Reddy GeedipallyAbstract:Crash data can often be characterized by over-dispersion, heavy (long) tail and many observations with the value zero. Over the last few years, a small number of researchers have started developing and applying novel and innovative multi-parameter Models to analyze such data. These multi-parameter Models have been proposed for overcoming the limitations of the traditional negative binomial (NB) Model, which cannot handle this kind of data efficiently. The research documented in this paper continues the work related to multi-parameter Models. The objective of this paper is to document the development and application of a flexible NB Generalized Linear Model with randomly distributed mixed effects characterized by the Dirichlet process (NB-DP) to Model crash data. The objective of the study was accomplished using two datasets. The new Model was compared to the NB and the recently introduced Model based on the mixture of the NB and Lindley (NB-L) distributions. Overall, the research study shows that the NB-DP Model offers a better performance than the NB Model once data are over-dispersed and have a heavy tail. The NB-DP performed better than the NB-L when the dataset has a heavy tail, but a smaller percentage of zeros. However, both Models performed similarly when the dataset contained a large amount of zeros. In addition to a greater flexibility, the NB-DP provides a clustering by-product that allows the safety analyst to better understand the characteristics of the data, such as the identification of outliers and sources of dispersion.
-
the negative binomial lindley Generalized Linear Model characteristics and application using crash data
Accident Analysis & Prevention, 2012Co-Authors: Srinivas Reddy Geedipally, Dominique Lord, Soma S DhavalaAbstract:There has been a considerable amount of work devoted by transportation safety analysts to the development and application of new and innovative Models for analyzing crash data. One important characteristic about crash data that has been documented in the literature is related to datasets that contained a large amount of zeros and a long or heavy tail (which creates highly dispersed data). For such datasets, the number of sites where no crash is observed is so large that traditional distributions and regression Models, such as the Poisson and Poisson-gamma or negative binomial (NB) Models cannot be used efficiently. To overcome this problem, the NB-Lindley (NB-L) distribution has recently been introduced for analyzing count data that are characterized by excess zeros. The objective of this paper is to document the application of a NB Generalized Linear Model with Lindley mixed effects (NB-L GLM) for analyzing traffic crash data. The study objective was accomplished using simulated and observed datasets. The simulated dataset was used to show the general performance of the Model. The Model was then applied to two datasets based on observed data. One of the dataset was characterized by a large amount of zeros. The NB-L GLM was compared with the NB and zero-inflated Models. Overall, the research study shows that the NB-L GLM not only offers superior performance over the NB and zero-inflated Models when datasets are characterized by a large number of zeros and a long tail, but also when the crash dataset is highly dispersed.
Richard J Samworth - One of the best experts on this subject based on the ideXlab platform.
-
goodness of fit testing in high dimensional Generalized Linear Models
Journal of The Royal Statistical Society Series B-statistical Methodology, 2020Co-Authors: Jana Jankova, Rajen D Shah, Peter Buhlmann, Richard J SamworthAbstract:We propose a family of tests to assess the goodness of fit of a high dimensional Generalized Linear Model. Our framework is flexible and may be used to construct an omnibus test or directed against testing specific non‐Linearities and interaction effects, or for testing the significance of groups of variables. The methodology is based on extracting left‐over signal in the residuals from an initial fit of a Generalized Linear Model. This can be achieved by predicting this signal from the residuals by using modern powerful regression or machine learning methods such as random forests or boosted trees. Under the null hypothesis that the Generalized Linear Model is correct, no signal is left in the residuals and our test statistic has a Gaussian limiting distribution, translating to asymptotic control of type I error. Under a local alternative, we establish a guarantee on the power of the test. We illustrate the effectiveness of the methodology on simulated and real data examples by testing goodness of fit in logistic regression Models. Software implementing the methodology is available in the R package GRPtests.
-
goodness of fit testing in high dimensional Generalized Linear Models
arXiv: Methodology, 2019Co-Authors: Jana Jankova, Rajen D Shah, Peter Buhlmann, Richard J SamworthAbstract:We propose a family of tests to assess the goodness-of-fit of a high-dimensional Generalized Linear Model. Our framework is flexible and may be used to construct an omnibus test or directed against testing specific non-Linearities and interaction effects, or for testing the significance of groups of variables. The methodology is based on extracting left-over signal in the residuals from an initial fit of a Generalized Linear Model. This can be achieved by predicting this signal from the residuals using modern flexible regression or machine learning methods such as random forests or boosted trees. Under the null hypothesis that the Generalized Linear Model is correct, no signal is left in the residuals and our test statistic has a Gaussian limiting distribution, translating to asymptotic control of type I error. Under a local alternative, we establish a guarantee on the power of the test. We illustrate the effectiveness of the methodology on simulated and real data examples by testing goodness-of-fit in logistic regression Models. Software implementing the methodology is available in the R package `GRPtests'.