The Experts below are selected from a list of 56673 Experts worldwide ranked by ideXlab platform

Itshak Lapidot - One of the best experts on this subject based on the ideXlab platform.

  • Speakers Clustering with stochastic VQ and Clustering Quality estimator
    2018 IEEE International Conference on the Science of Electrical Engineering in Israel (ICSEE), 2018
    Co-Authors: Yishai Cohen, Itshak Lapidot
    Abstract:

    Short segments speaker Clustering has significant importance both for diarization and applications such as short push-to-tatk (PTT) segments Clustering. In this paper we present a new way to cluster speech segments by applying a stochastic vector quantization (VQ) with a cosine metric together with a speaker Clustering Quality estimator based on logistic regression. The VQ is performed on codebooks of different sizes, and the choice of the best Clustering result is estimated using logistic regression. The algorithm is tested on a large range of speakers, between 2 to 60. The results are compared to those of the mean-shift Clustering method, which was already tested for this task several times. The results are a bit below those of the cosine similarity measure-based mean-shift Clustering. The advantage is in the run-time which is approximately 10 times faster.

  • Robust speaker Clustering Quality estimation
    2018 IEEE International Conference on the Science of Electrical Engineering in Israel (ICSEE), 2018
    Co-Authors: Yishai Cohen, Itshak Lapidot
    Abstract:

    This paper focuses on estimating the Quality of a Clustering process. In our case - the task is to cluster short speech segments that belong to different speakers. Moreover, speaker Clustering Quality may be well estimated on several Clustering approaches if they all based on the same features. This is very important, as it allows us to use the same Quality estimation system without retraining, and achieve reasonable results even when the Clustering method is changed. We predict the system’s Quality by applying a logistic regression estimator on a several statistical parameters of the Clustering. In this paper, mean-shift Clustering with either cosine or probabilistic linear discriminant analysis (PLDA) score as similarity measure, and stochastic vector quantization (VQ) with cosine distance were applied in order to cluster the short speaker segments represented by i-vectors. The Quality of the Clustering is measured using the average cluster purity (ACP), average speaker purity (ASP) and K. We show that these measures can be estimated fairly well by applying logistic regression based on various Clustering statistics that calculated once Clustering is over. These statistical parameters are used as a feature vector representing the Clustering.

  • Speaker Clustering Quality estimation with logistic regression
    Computer Speech & Language, 1
    Co-Authors: Yishai Cohen, Itshak Lapidot
    Abstract:

    Abstract This paper focuses on estimating the Quality of a Clustering process. The task is to cluster short speech segments that belong to different speakers. A variety of statistical parameters are estimated from the output of the Clustering process. These parameters are used to train a logistic regression to serve as a Clustering Quality estimation system. In this paper, mean-shift Clustering with either a cosine distance or probabilistic linear discriminant analysis (PLDA) score as the similarity measure, as well as stochastic vector quantization (VQ) with cosine distance, are applied in order to cluster the short speaker segments, which are represented by i-vectors. The Quality of the Clustering is measured using the average cluster purity (ACP), average speaker purity (ASP) and K, which is the geometric mean of ASP and ACP. We show that these measures can be estimated fairly well by applying logistic regression. Moreover, Clustering Quality may be well estimated even if the logistic regression was trained using parameters derived from a different Clustering algorithm. This is very important, as it allows the use of a single Quality estimation system, without the need for retraining when the Clustering method is changed. Additionally, we showed how the Clustering Quality estimator could be served as an estimator of the number of clusters. For VQ-based Clustering the number of clusters has to be predefined. We perform the Clustering with different number of clusters. The best number of clusters is estimated as the Clustering that achieved the higher estimation of the K value. We will show that this approach estimate the best number of clusters accurately.

Louis Massey - One of the best experts on this subject based on the ideXlab platform.

  • on the Quality of art1 text Clustering
    International Joint Conference on Neural Network, 2003
    Co-Authors: Louis Massey
    Abstract:

    There is a large and continually growing quantity of electronic text available, which contain essential human and organization knowledge. An important research endeavor is to study and develop better ways to access this knowledge. Text Clustering is a popular approach to automatically organize textual document collections by topics to help users find the information they need. Adaptive Resonance Theory (ART) neural networks possess several interesting properties that make them appealing in the area of text Clustering. Although ART has been used in several research works as a text Clustering tool, the level of Quality of the resulting document clusters has not been clearly established yet. In this paper, we present experimental results with binary ART that address this issue by determining how close Clustering Quality is to an upper bound on Clustering Quality.

Tuanh Hoang Nguyen - One of the best experts on this subject based on the ideXlab platform.

  • using textual semantic similarity to improve Clustering Quality of web video search results
    Knowledge and Systems Engineering, 2015
    Co-Authors: Phuc Quang Nguyen, Anhthu Nguyenthi, Thanh Duc Ngo, Tuanh Hoang Nguyen
    Abstract:

    Clustering Web video search results is to help users locating videos of interest in more effective manner. To cluster returned videos, existing works proposed to use textual and visual similarity of videos. However, one of their limitations is that semantic similarity of textual metadata was not considered. Meanwhile, metadata of videos are usually annotated by users with words of high semantic level. This paper introduces a thesaurus based approach to estimate textual semantic similarity of metadata for Clustering Web video search results. Experiments were conducted on a set of real-world videos crawled from the Internet. The experimental results demonstrated that using semantic similarity of textual metadata in the combination with visual similarity significantly improves Clustering Quality.

  • KSE - Using Textual Semantic Similarity to Improve Clustering Quality of Web Video Search Results
    2015 Seventh International Conference on Knowledge and Systems Engineering (KSE), 2015
    Co-Authors: Phuc Quang Nguyen, Thanh Duc Ngo, Anh-thu Nguyen-thi, Tuanh Hoang Nguyen
    Abstract:

    Clustering Web video search results is to help users locating videos of interest in more effective manner. To cluster returned videos, existing works proposed to use textual and visual similarity of videos. However, one of their limitations is that semantic similarity of textual metadata was not considered. Meanwhile, metadata of videos are usually annotated by users with words of high semantic level. This paper introduces a thesaurus based approach to estimate textual semantic similarity of metadata for Clustering Web video search results. Experiments were conducted on a set of real-world videos crawled from the Internet. The experimental results demonstrated that using semantic similarity of textual metadata in the combination with visual similarity significantly improves Clustering Quality.

Jean-charles Lamirel - One of the best experts on this subject based on the ideXlab platform.

  • New efficient Clustering Quality indexes
    2016
    Co-Authors: Jean-charles Lamirel, Nicolas Dugué, Pascal Cuxac
    Abstract:

    This paper deals with a major challenge in Clustering that is optimal model selection. It presents new efficient Clustering Quality indexes relying on feature maximization, which is an alternative measure to usual distributional measures relying on entropy, Chi-square metric or vector-based measures such as Euclidean distance or correlation distance. First Experiments compare the behavior of these new indexes with usual cluster Quality indexes based on Euclidean distance on different kinds of test datasets for which ground truth is available. This comparison clearly highlights altogether the superior accuracy and stability of the new method on these datasets, its efficiency from low to high dimensional range and its tolerance to noise. Further experiments are then conducted on " real life " textual data extracted from a multisource bibliographic database for which ground truth is unknown. These experiments show that the accuracy and stability of these new indexes allow to deal efficiently with diachronic analysis, when other indexes do not fit the requirements for this task.

  • WSOM - Reliable Clustering Quality Estimation from Low to High Dimensional Data
    Advances in Self-Organizing Maps and Learning Vector Quantization, 2016
    Co-Authors: Jean-charles Lamirel
    Abstract:

    This paper presents new cluster Quality indexes which can be efficiently applied for a low-to-high dimensional range of data and which are tolerant to noise. These indexes relies on feature maximization, which is an alternative measure to usual distributional measures relying on entropy or on Chi-square metric or vector-based measures such as Euclidean distance or correlation distance. Experiments compare the behavior of these new indexes with usual cluster Quality indexes based on Euclidean distance on different kinds of test datasets for which ground truth is available. This comparison clearly highlights the superior accuracy and stability of the new method.

  • IJCNN - New efficient Clustering Quality indexes
    2016 International Joint Conference on Neural Networks (IJCNN), 2016
    Co-Authors: Jean-charles Lamirel, Nicolas Dugué, Pascal Cuxac
    Abstract:

    This paper deals with a major challenge in Clustering that is optimal model selection. It presents new efficient Clustering Quality indexes relying on feature maximization, which is an alternative measure to usual distributional measures relying on entropy, Chi-square metric or vector-based measures such as Euclidean distance or correlation distance. First Experiments compare the behavior of these new indexes with usual cluster Quality indexes based on Euclidean distance on different kinds of test datasets for which ground truth is available. This comparison clearly highlights altogether the superior accuracy and stability of the new method on these datasets, its efficiency from low to high dimensional range and its tolerance to noise. Further experiments are then conducted on “real life” textual data extracted from a multisource bibliographic database for which ground truth is unknown. These experiments show that the accuracy and stability of these new indexes allow to deal efficiently with diachronic analysis, when other indexes do not fit the requirements for this task.

  • PAKDD Workshops - Feature Maximization Based Clustering Quality Evaluation: A Promising Approach
    Lecture Notes in Computer Science, 2015
    Co-Authors: Jean-charles Lamirel, Shadi Al Shehabi
    Abstract:

    Feature maximization is an alternative measure, as compared to usual distributional measures relying on entropy or on Chi-square metric or vector-based measures, like Euclidean distance or correlation distance. One of the key advantages of this measure is that it is operational in an incremental mode both on Clustering and on traditional classification. In the classification framework, it does not presents the limitations of the aforementioned measures in the case of the processing of highly unbalanced, heterogeneous and highly multidimensional data. We present a new application of this measure in the Clustering context for setting up new cluster Quality indexes whose efficiency ranges for low to high dimensional data and that are tolerant to noise. We compare the behaviour of these new indexes with usual cluster Quality indexes based on Euclidean distance on different kinds of test datasets for which ground truth is available. Proposed comparison clearly highlights the superior accuracy and stability of the new method.

  • PAKDD Workshops - A new efficient and unbiased approach for Clustering Quality evaluation
    New Frontiers in Applied Data Mining, 2012
    Co-Authors: Jean-charles Lamirel, Pascal Cuxac, Raghvendra Mall, Ghada Safi
    Abstract:

    Traditional Quality indexes (Inertia, DB, …) are known to be method-dependent indexes that do not allow to properly estimate the Quality of the Clustering in several cases, as in that one of complex data, like textual data. We thus propose an alternative approach for Clustering Quality evaluation based on unsupervised measures of Recall, Precision and F-measure exploiting the descriptors of the data associated with the obtained clusters. Two categories of index are proposed, that are Macro and Micro indexes. This paper also focuses on the construction of a new cumulative Micro precision index that makes it possible to evaluate the overall Quality of a Clustering result while clearly distinguishing between homogeneous and heterogeneous, or degenerated results. The experimental comparison of the behavior of the classical indexes with our new approach is performed on a polythematic dataset of bibliographical references issued from the PASCAL database.

Yishai Cohen - One of the best experts on this subject based on the ideXlab platform.

  • Speakers Clustering with stochastic VQ and Clustering Quality estimator
    2018 IEEE International Conference on the Science of Electrical Engineering in Israel (ICSEE), 2018
    Co-Authors: Yishai Cohen, Itshak Lapidot
    Abstract:

    Short segments speaker Clustering has significant importance both for diarization and applications such as short push-to-tatk (PTT) segments Clustering. In this paper we present a new way to cluster speech segments by applying a stochastic vector quantization (VQ) with a cosine metric together with a speaker Clustering Quality estimator based on logistic regression. The VQ is performed on codebooks of different sizes, and the choice of the best Clustering result is estimated using logistic regression. The algorithm is tested on a large range of speakers, between 2 to 60. The results are compared to those of the mean-shift Clustering method, which was already tested for this task several times. The results are a bit below those of the cosine similarity measure-based mean-shift Clustering. The advantage is in the run-time which is approximately 10 times faster.

  • Robust speaker Clustering Quality estimation
    2018 IEEE International Conference on the Science of Electrical Engineering in Israel (ICSEE), 2018
    Co-Authors: Yishai Cohen, Itshak Lapidot
    Abstract:

    This paper focuses on estimating the Quality of a Clustering process. In our case - the task is to cluster short speech segments that belong to different speakers. Moreover, speaker Clustering Quality may be well estimated on several Clustering approaches if they all based on the same features. This is very important, as it allows us to use the same Quality estimation system without retraining, and achieve reasonable results even when the Clustering method is changed. We predict the system’s Quality by applying a logistic regression estimator on a several statistical parameters of the Clustering. In this paper, mean-shift Clustering with either cosine or probabilistic linear discriminant analysis (PLDA) score as similarity measure, and stochastic vector quantization (VQ) with cosine distance were applied in order to cluster the short speaker segments represented by i-vectors. The Quality of the Clustering is measured using the average cluster purity (ACP), average speaker purity (ASP) and K. We show that these measures can be estimated fairly well by applying logistic regression based on various Clustering statistics that calculated once Clustering is over. These statistical parameters are used as a feature vector representing the Clustering.

  • Speaker Clustering Quality estimation with logistic regression
    Computer Speech & Language, 1
    Co-Authors: Yishai Cohen, Itshak Lapidot
    Abstract:

    Abstract This paper focuses on estimating the Quality of a Clustering process. The task is to cluster short speech segments that belong to different speakers. A variety of statistical parameters are estimated from the output of the Clustering process. These parameters are used to train a logistic regression to serve as a Clustering Quality estimation system. In this paper, mean-shift Clustering with either a cosine distance or probabilistic linear discriminant analysis (PLDA) score as the similarity measure, as well as stochastic vector quantization (VQ) with cosine distance, are applied in order to cluster the short speaker segments, which are represented by i-vectors. The Quality of the Clustering is measured using the average cluster purity (ACP), average speaker purity (ASP) and K, which is the geometric mean of ASP and ACP. We show that these measures can be estimated fairly well by applying logistic regression. Moreover, Clustering Quality may be well estimated even if the logistic regression was trained using parameters derived from a different Clustering algorithm. This is very important, as it allows the use of a single Quality estimation system, without the need for retraining when the Clustering method is changed. Additionally, we showed how the Clustering Quality estimator could be served as an estimator of the number of clusters. For VQ-based Clustering the number of clusters has to be predefined. We perform the Clustering with different number of clusters. The best number of clusters is estimated as the Clustering that achieved the higher estimation of the K value. We will show that this approach estimate the best number of clusters accurately.