The Experts below are selected from a list of 18273 Experts worldwide ranked by ideXlab platform
Kenji Yamanishi - One of the best experts on this subject based on the ideXlab platform.
-
KDD - Tracking dynamics of topic trends using a Finite Mixture Model
Proceedings of the 2004 ACM SIGKDD international conference on Knowledge discovery and data mining - KDD '04, 2004Co-Authors: Satoshi Morinaga, Kenji YamanishiAbstract:In a wide range of business areas dealing with text data streams, including CRM, knowledge management, and Web monitoring services, it is an important issue to discover topic trends and analyze their dynamics in real-time. Specifically we consider the following three tasks in topic trend analysis: 1)Topic Structure Identification; identifying what kinds of main topics exist and how important they are, 2)Topic Emergence Detection; detecting the emergence of a new topic and recognizing how it grows, 3)Topic Characterization; identifying the characteristics for each of main topics. For real topic analysis systems, we may require that these three tasks be performed in an on-line fashion rather than in a retrospective way, and be dealt with in a single framework. This paper proposes a new topic analysis framework which satisfies this requirement from a unifying viewpoint that a topic structure is Modeled using a Finite Mixture Model and that any change of a topic trend is tracked by learning the Finite Mixture Model dynamically. In this framework we propose the usage of a time-stamp based discounting learning algorithm in order to realize real-time topic structure identification. This enables tracking the topic structure adaptively by forgetting out-of-date statistics. Further we apply the theory of dynamic Model selection to detecting changes of main components in the Finite Mixture Model in order to realize topic emergence detection. We demonstrate the effectiveness of our framework using real data collected at a help desk to show that we are able to track dynamics of topic trends in a timely fashion.
-
Tracking dynamics of topic trends using a Finite Mixture Model
2004Co-Authors: Satoshi Morinaga, Kenji YamanishiAbstract:In a wide range of business areas dealing with text data streams, including CRM, knowledge management, and Web monitoring services, it is an important issue to discover topic trends and analyze their dynamics in real-time. Specifically we consider the following three tasks in topic trend analysis: 1)Topic Structure Identification; identifying what kinds of main topics exist and how important they are, 2)Topic Emergence Detection; detecting the emergence of a new topic and recognizing how it grows, 3)Topic Characterization; identifying the characteristics for each of main topics. For real topic analysis systems, we may require that these three tasks be performed in an on-line fashion rather than in a retrospective way, and be dealt with in a single framework. This paper proposes a new topic analysis framework which satisfies this requirement from a unifying viewpoint that a topic structure is Modeled using a Finite Mixture Model and that any change of a topic trend is tracked by learning the Finite Mixture Model dynamically. In this framework we propose the usage of a time-stamp based discounting learning algorithm in order to realize real-time topic structure identification. This enables tracking the topic structure adaptively by forgetting out-of-date statistics. Further we apply the theory of dynamic Model selection to detecting changes of main components in the Finite Mixture Model in order to realize topic emergence detection. We demonstrate the effectiveness of our framework using real data collected at a help desk to show that we are able to track dynamics of topic trends in a timely fashion.
-
Topic analysis using a Finite Mixture Model
Information Processing & Management, 2003Co-Authors: Kenji YamanishiAbstract:Addressed here is the issue of 'topic analysis' which is used to determine a text's topic structure, a representation indicating what topics are included in a text and how those topics change within the text. Topic analysis consists of two main tasks: topic identification and text segmentation. While topic analysis would be extremely useful in a variety of text processing applications, no previous study has so far sufficiently addressed it. A statistical learning approach to the issue is proposed in this paper. More specifically, topics here are represented by means of word clusters, and a Finite Mixture Model, referred to as a stochastic topic Model (STM), is employed to represent a word distribution within a text. In topic analysis, a given text is segmented by detecting significant differences between STMs, and topics are identified by means of estimation of STMs. Experimental results indicate that the proposed method significantly outperforms methods that combine existing techniques.
-
topic analysis using a Finite Mixture Model
Empirical Methods in Natural Language Processing, 2000Co-Authors: Kenji YamanishiAbstract:We address the issue of 'topic analysis,' by which is determined a text's topic structure, which indicates what topics are included in a text, and how topics change within the text. We propose a novel approach to this issue, one based on statistical Modeling and learning. We represent topics by means of word clusters, and employ a Finite Mixture Model to represent a word distribution within a text. Our experimental results indicate that our method significantly outperforms a method that combines existing techniques.
-
EMNLP - Topic Analysis Using a Finite Mixture Model
Proceedings of the 2000 Joint SIGDAT conference on Empirical methods in natural language processing and very large corpora held in conjunction with th, 2000Co-Authors: Kenji YamanishiAbstract:We address the issue of 'topic analysis,' by which is determined a text's topic structure, which indicates what topics are included in a text, and how topics change within the text. We propose a novel approach to this issue, one based on statistical Modeling and learning. We represent topics by means of word clusters, and employ a Finite Mixture Model to represent a word distribution within a text. Our experimental results indicate that our method significantly outperforms a method that combines existing techniques.
Pravin K. Trivedi - One of the best experts on this subject based on the ideXlab platform.
-
equity in swedish health care reconsidered new results based on the Finite Mixture Model
Health Economics, 2001Co-Authors: Ulf Gerdtham, Pravin K. TrivediAbstract:This paper reconsiders the equity issue in Swedish health care utilization previously analysed by Gerdtham (Health Econ 1997; 6: 303-319) within the framework of the standard two-part Model. Departing from the user/non-user distinction, we use the more flexible framework of the Finite Mixture Model that distinguishes between frequent/infrequent users. Our results indicate that the support for the inequity hypothesis reported by Gerdtham is sensitive to Model specification and the way standard errors of coefficients are estimated. The new framework offers an alternative perspective on the magnitude of the income-related difference in health care utilization.
-
Equity in Swedish Health Care Reconsidered: New Results Based on the Finite Mixture Model
SSRN Electronic Journal, 2000Co-Authors: Ulf Gerdtham, Pravin K. TrivediAbstract:This paper reconsiders the equity issue in Swedish health care utilisation previously analyzed by Gerdtham (Health Economics 6, 303-319, 1997) within the framework of the standard two-part Model. Departing from the user/nonuser distinction, we use the more flexible framework of the Finite Mixture Model that distinguishes between frequent/infrequent users. Our empirical results indicate that the Finite Mixture Model fits the data better than the two-part Model. The results indicate that income increases use among infrequent users in both physician and hospital care but that there is no such income effect among frequent users.
Ulf Gerdtham - One of the best experts on this subject based on the ideXlab platform.
-
equity in swedish health care reconsidered new results based on the Finite Mixture Model
Health Economics, 2001Co-Authors: Ulf Gerdtham, Pravin K. TrivediAbstract:This paper reconsiders the equity issue in Swedish health care utilization previously analysed by Gerdtham (Health Econ 1997; 6: 303-319) within the framework of the standard two-part Model. Departing from the user/non-user distinction, we use the more flexible framework of the Finite Mixture Model that distinguishes between frequent/infrequent users. Our results indicate that the support for the inequity hypothesis reported by Gerdtham is sensitive to Model specification and the way standard errors of coefficients are estimated. The new framework offers an alternative perspective on the magnitude of the income-related difference in health care utilization.
-
Equity in Swedish Health Care Reconsidered: New Results Based on the Finite Mixture Model
SSRN Electronic Journal, 2000Co-Authors: Ulf Gerdtham, Pravin K. TrivediAbstract:This paper reconsiders the equity issue in Swedish health care utilisation previously analyzed by Gerdtham (Health Economics 6, 303-319, 1997) within the framework of the standard two-part Model. Departing from the user/nonuser distinction, we use the more flexible framework of the Finite Mixture Model that distinguishes between frequent/infrequent users. Our empirical results indicate that the Finite Mixture Model fits the data better than the two-part Model. The results indicate that income increases use among infrequent users in both physician and hospital care but that there is no such income effect among frequent users.
Hani S Mahmassani - One of the best experts on this subject based on the ideXlab platform.
-
a Finite Mixture Model of vehicle to vehicle and day to day variability of traffic network travel times
Transportation Research Part C-emerging Technologies, 2014Co-Authors: Jiwon Kim, Hani S MahmassaniAbstract:This study proposes an approach to Modeling the effects of daily roadway conditions on travel time variability using a Finite Mixture Model based on the Gamma–Gamma (GG) distribution. The GG distribution is a compound distribution derived from the product of two Gamma random variates, which represent vehicle-to-vehicle and day-to-day variability, respectively. It provides a systematic way of investigating different variability dimensions reflected in travel time data. To identify the underlying distribution of each type of variability, this study first decomposes a Mixture of Gamma–Gamma Models into two separate Gamma Mixture Modeling problems and estimates the respective parameters using the Expectation–Maximization (EM) algorithm. The proposed methodology is demonstrated using simulated vehicle trajectories produced under daily scenarios constructed from historical weather and accident data. The parameter estimation results suggest that day-to-day variability exhibits clear heterogeneity under different weather conditions: clear versus rainy or snowy days, whereas the same weather conditions have little impact on vehicle-to-vehicle variability. Next, a two-component Gamma–Gamma Mixture Model is specified. The results of the distribution fitting show that the Mixture Model provides better fits to travel delay observations than the standard (one-component) Gamma–Gamma Model. The proposed method, the application of the compound Gamma distribution combined with a Mixture Modeling approach, provides a powerful and flexible tool to capture not only different types of variability—vehicle-to-vehicle and day-to-day variability—but also the unobserved heterogeneity within these variability types, thereby allowing the Modeling of the underlying distributions of individual travel delays across different days with varying roadway disruption levels in a more effective and systematic way.
-
a Finite Mixture Model of vehicle to vehicle and day to day variability of traffic network travel times
Transportation Research Board 93rd Annual MeetingTransportation Research Board, 2014Co-Authors: Jiwon Kim, Hani S MahmassaniAbstract:This study proposes an approach to analyze the effects of daily roadway conditions on travel time variability using a Finite Mixture Model based on the Gamma-Gamma (GG) distribution. The GG distribution is a compound distribution derived from the product of two Gamma random variates, which represent vehicle-to-vehicle and day-to-day variability, respectively. It provides a systematic way of investigating different variability dimensions reflected in travel time data. To identify the underlying distribution of each type of variability, this study first decomposes a Mixture of GG Models into two separate Gamma Mixture Modeling problems and estimates the respective parameters using the Expectation-Maximization algorithm. The proposed methodology is demonstrated using simulated vehicle trajectories produced under daily scenarios constructed from historical weather and accident data. The parameter estimation results suggest that day-to-day variability exhibits clear heterogeneity under different weather conditions: clear versus rainy or snowy days, whereas the same weather conditions have little impact on vehicle-to-vehicle variability. Next, a two-component GG Mixture Model is specified. The distribution fitting results show that the Mixture Model provides better fits to travel delay observations than the standard (one-component) GG Model. The proposed method, the application of the compound Gamma distribution combined with a Mixture Modeling approach, provides a powerful and flexible tool to capture not only different types of variability—vehicle-to-vehicle and day-to-day variability—but also the unobserved heterogeneity within these variability types, thereby allowing a better understanding of the mechanisms causing travel time unreliability in a traffic network.
Marco Alfò - One of the best experts on this subject based on the ideXlab platform.
-
Finite Mixture Model-based classification of a complex vegetation system
Vegetation Classification and Survey, 2020Co-Authors: Fabio Attorre, Vito Emanuele Cambria, Emiliano Agrillo, Nicola Alessi, Marco Alfò, Michele De Sanctis, Luca Malatesta, Tommaso Sitzia, Riccardo Guarino, Corrado MarcenòAbstract:Aim: To propose a Finite Mixture Model (FMM) as an additional approach for classifying large datasets of georeferenced vegetation plots from complex vegetation systems. Study area: The Italian peninsula including the two main islands (Sicily and Sardinia), but excluding the Alps and the Po plain. Methods: We used a database of 5,593 georeferenced plots and 1,586 vascular species of forest vegetation, created in TURBOVEG by storing published and unpublished phytosociological plots collected over the last 30 years. The plots were classified according to species composition and environmental variables using a FMM. Classification results were compared with those obtained by TWINSPAN algorithm. Groups were characterized in terms of ecological parameters, dominant and diagnostic species using the fidelity coefficient. Interpretation of resulting forest vegetation types was supported by a predictive map, produced using discriminant functions on environmental predictors, and by a non‐metric multidimensional scaling ordination. Results: FMM clustering obtained 24 groups that were compared with those from TWINSPAN, and similarities were found only at a higher classification level corresponding to the main orders of the Italian broadleaf forest vegetation: Fagetalia sylvaticae, Carpinetalia betuli, Quercetalia pubescenti-petraeae and Quercetalia ilicis. At lower syntaxonomic level, these 24 groups were referred to alliances and sub-alliances. Conclusions: Despite a greater computational complexity, FMM appears to be an effective alternative to the traditional classification methods through the incorporation of Modelling in the classificatory process. This allows classification of both the co-occurrence of species and environmental factors so that groups are identified not only on their species composition, as in the case of TWINSPAN, but also on their specific environmental niche. Taxonomic reference: Conti et al. (2005). Abbreviations: CLM = Community-level Models; FMM = Finite Mixture Model; NMDS = non‐metric multidimensional scaling.
-
a bidimensional Finite Mixture Model for longitudinal data subject to dropout
Statistics in Medicine, 2018Co-Authors: Alessandra Spagnoli, Maria Francesca Marino, Marco AlfòAbstract:In longitudinal studies, subjects may be lost to follow up and, thus, present incomplete response sequences. When the mechanism underlying the dropout is nonignorable, we need to account for dependence between the longitudinal and the dropout process. We propose to Model such a dependence through discrete latent effects, which are outcome-specific and account for heterogeneity in the univariate profiles. Dependence between profiles is introduced by using a probability matrix to describe the corresponding joint distribution. In this way, we separately Model dependence within each outcome and dependence between outcomes. The major feature of this proposal, when compared with standard Finite Mixture Models, is that it allows the nonignorable dropout Model to properly nest its ignorable counterpart. We also discuss the use of an index of (local) sensitivity to nonignorability to investigate the effects that assumptions about the dropout process may have on Model parameter estimates. The proposal is illustrated via the analysis of data from a longitudinal study on the dynamics of cognitive functioning in the elderly.
-
Classifying and Mapping Potential Distribution of Forest Types Using a Finite Mixture Model
Folia Geobotanica, 2014Co-Authors: Fabio Attorre, Marco Alfò, Michele De Sanctis, Fabio Francesconi, Francesca Martella, Roberto Valenti, Marcello VitaleAbstract:The present paper presents the application of a Finite Mixture Model (FMM) to analyze spatially explicit data on forest composition and environmental variables to produce a high-resolution map of their current potential distribution. FMM provides a convenient yet formal setting for Model-based clustering. Within this framework, forest data are assumed to come from an underlying FMM, where each Mixture component corresponds to a cluster and each cluster is characterized by a different composition of tree species. An important extension of this Model is based on including a set of covariates to predict class membership. These covariates can be climatic and topographical parameters as well as geographical coordinates and the class membership of neighbouring plots. FMM was applied to a national forest inventory of Italy consisting of 6,714 plots with a measure of abundance for 27 tree species. In this way, a map of potential forest types was produced. The limitations and usefulness of the proposed Modelling approach were analyzed and discussed, comparing the results with an independently derived expert map.
-
a Finite Mixture Model for multivariate counts under endogenous selectivity
Statistics and Computing, 2011Co-Authors: Marco Alfò, Antonello Maruotti, Giovanni TrovatoAbstract:We describe a selection Model for multivariate counts, where association between the primary outcomes and the endogenous selection source is Modeled through outcome-specific latent effects which are assumed to be dependent across equations. Parametric specifications of this Model already exist in the literature; in this paper, we show how Model parameters can be estimated in a Finite Mixture context. This approach helps us to consider overdispersed counts, while allowing for multivariate association and endogeneity of the selection variable. In this context, attention is focused both on bias in estimated effects when exogeneity of selection (treatment) variable is assumed, as well as on consistent estimation of the association between the random effects in the primary and in the treatment effect Models, when the latter is assumed endogeneous. The Model behavior is investigated through a large scale simulation experiment. An empirical example on health care utilization data is provided.
-
a Finite Mixture Model for image segmentation
Statistics and Computing, 2008Co-Authors: Marco Alfò, Luciano Nieddu, Donatella VicariAbstract:In this paper, we propose a Model for image segmentation based on a Finite Mixture of Gaussian distributions. For each pixel of the image, prior probabilities of class memberships are specified through a Gibbs distribution, where association between labels of adjacent pixels is Modeled by a class-specific term allowing for different interaction strengths across classes. We show how Model parameters can be estimated in a maximum likelihood framework using Mean Field theory. Experimental performance on perturbed phantom and on real benchmark images shows that the proposed method performs well in a wide variety of empirical situations.