The Experts below are selected from a list of 7512 Experts worldwide ranked by ideXlab platform
Setsuo Arikawa - One of the best experts on this subject based on the ideXlab platform.
-
CPM - Efficient Discovery of Proximity Patterns with Suffix Arrays
2001Co-Authors: Hiroki Arimura, Hiroki Asaka, Hiroshi Sakamoto, Setsuo ArikawaAbstract:We describe an efficient implementation of a Text Mining Algorithm for discovering a class of simple string patterns. With an index structure, called the virtual suffix tree, for pattern discovery built on the top of the suffix array, the resulting Algorithm is simple and fast in practice compared with the previous implementation with the suffix tree.
-
Efficient Discovery of Proximity Patterns with Suffix Arrays (Extended Abstract)
Combinatorial Pattern Matching, 2001Co-Authors: Hiroki Arimura, Hiroki Asaka, Hiroshi Sakamoto, Setsuo ArikawaAbstract:We describe an efficient implementation of a Text Mining Algorithm for discovering a class of simple string patterns. With an index structure, called the virtual suffix tree, for pattern discovery built on the top of the suffix array, the resulting Algorithm is simple and fast in practice compared with the previous implementation with the suffix tree.
-
Efficient discovery of proximity patterns with suffix arrays
Lecture Notes in Computer Science, 2001Co-Authors: Hiroki Arimura, Hiroki Asaka, Hiroshi Sakamoto, Setsuo ArikawaAbstract:We describe an efficient implementation of a Text Mining Algorithm for discovering a class of simple string patterns. With an index structure, called the virtual suffix tree, for pattern discovery built on the top of the suffix array, the resulting Algorithm is simple and fast in practice compared with the previous implementation with the suffix tree.
Kenneth W. Kizer - One of the best experts on this subject based on the ideXlab platform.
-
A Text-Mining approach to obtain detailed treatment information from free-Text fields in population-based cancer registries: A study of non-small cell lung cancer in California
PloS one, 2019Co-Authors: Frances B. Maguire, Cyllene R. Morris, Arti Parikh-patel, Rosemary D. Cress, Theresa H. M. Keegan, Patrick S. Lin, Kenneth W. KizerAbstract:Author(s): Maguire, Frances B; Morris, Cyllene R; Parikh-Patel, Arti; Cress, Rosemary D; Keegan, Theresa H M; Li, Chin-Shang; Lin, Patrick S; Kizer, Kenneth W | Abstract: Population-based cancer registries have treatment information for all patients making them an excellent resource for population-level monitoring. However, specific treatment details, such as drug names, are contained in a free-Text format that is difficult to process and summarize. We assessed the accuracy and efficiency of a Text-Mining Algorithm to identify systemic treatments for lung cancer from free-Text fields in the California Cancer Registry.The Algorithm used Perl regular expressions in SAS 9.4 to search for treatments in 24,845 free-Text records associated with 17,310 patients in California diagnosed with stage IV non-small cell lung cancer between 2012 and 2014. Our Algorithm categorized treatments into six groups that align with National Comprehensive Cancer Network guidelines. We compared results to a manual review (gold standard) of the same records.Percent agreement ranged from 91.1% to 99.4%. Ranges for other measures were 0.71-0.92 (Kappa), 74.3%-97.3% (sensitivity), 92.4%-99.8% (specificity), 60.4%-96.4% (positive predictive value), and 92.9%-99.9% (negative predictive value). The Text-Mining Algorithm used one-sixth of the time required for manual review.SAS-based Text Mining of free-Text data can accurately detect systemic treatments administered to patients and save considerable time compared to manual review, maximizing the utility of the extant information in population-based cancer registries for comparative effectiveness research.
-
Agreement of treatment between the SAS Text Mining Algorithm and manual review among stage IV non-small cell lung cancer patients (n = 17, 310), 2012–2014, California.
2019Co-Authors: Frances B. Maguire, Cyllene R. Morris, Arti Parikh-patel, Rosemary D. Cress, Theresa H. M. Keegan, Patrick S. Lin, Kenneth W. KizerAbstract:Agreement of treatment between the SAS Text Mining Algorithm and manual review among stage IV non-small cell lung cancer patients (n = 17, 310), 2012–2014, California.
Wen Zou - One of the best experts on this subject based on the ideXlab platform.
-
Erratum to: A novel procedure on next generation sequencing data analysis using Text Mining Algorithm
BMC bioinformatics, 2016Co-Authors: Weizhong Zhao, James J. Chen, Roger Perkins, Yuping Wang, Zhichao Liu, Huixiao Hong, Weida Tong, Wen ZouAbstract:Erratum After publication of the original article [1] it was brought to our attention that the following was incorrectly placed under subheading ‘3. Classification analysis and comparison’ of subsection ‘Evaluation of topic modeling performance’ of the ‘Methods’ section: Topic model-derived clustering method [33] was applied, in which LDA was utilized as a feature reduction approach for cluster analysis. The LDAderived topics were considered as the new features of datasets. The sampletopic matrix (Fig. 1(f )) was treated as a new representation of the original dataset. Based on the sample-topic matrix (topic number was chosen as 5 and 30, respectively), conventional clustering Algorithms, such as k-means, was used for the clustering analysis. The number of clusters was set as 7 in the k-means method due to 7 different serotypes in the dataset. While in comparison, k-means Algorithm was also applied on VSM matrix using Hamming Distance similarities. For further comparison, due to the dimension reduction of topic modeling approach, the traditional tool of PCA was used to reduce features (Numbers of 2, 5, 10 and 30 were randomly selected as the reduced features, respectively) of VSM matrix followed by the k-means cluster analysis. Moreover, clustering by only LDA referred as “highest probable topic assignment” [33] (5 and 30 topics were used) was also used for comparison. In “highest probable topic assignment”, the LDA-derived topics were made as the clusters of the dataset. Then, each sample was assigned to the cluster (Topic) with the highest probability in the row of the sample-topic matrix. To interpret the clustering results obtained by the k-means Algorithm, samples in each cluster were labeled as the dominant serotype of the samples in the cluster. The
-
A novel procedure on next generation sequencing data analysis using Text Mining Algorithm
BMC bioinformatics, 2016Co-Authors: Weizhong Zhao, James J. Chen, Roger Perkins, Yuping Wang, Zhichao Liu, Huixiao Hong, Weida Tong, Wen ZouAbstract:Next-generation sequencing (NGS) technologies have provided researchers with vast possibilities in various biological and biomedical research areas. Efficient data Mining strategies are in high demand for large scale comparative and evolutional studies to be performed on the large amounts of data derived from NGS projects. Topic modeling is an active research field in machine learning and has been mainly used as an analytical tool to structure large Textual corpora for data Mining. We report a novel procedure to analyse NGS data using topic modeling. It consists of four major procedures: NGS data retrieval, preprocessing, topic modeling, and data Mining using Latent Dirichlet Allocation (LDA) topic outputs. The NGS data set of the Salmonella enterica strains were used as a case study to show the workflow of this procedure. The perplexity measurement of the topic numbers and the convergence efficiencies of Gibbs sampling were calculated and discussed for achieving the best result from the proposed procedure. The output topics by LDA Algorithms could be treated as features of Salmonella strains to accurately describe the genetic diversity of fliC gene in various serotypes. The results of a two-way hierarchical clustering and data matrix analysis on LDA-derived matrices successfully classified Salmonella serotypes based on the NGS data. The implementation of topic modeling in NGS data analysis procedure provides a new way to elucidate genetic information from NGS data, and identify the gene-phenotype relationships and biomarkers, especially in the era of biological and medical big data. The implementation of topic modeling in NGS data analysis provides a new way to elucidate genetic information from NGS data, and identify the gene-phenotype relationships and biomarkers, especially in the era of biological and medical big data.
Hiroki Arimura - One of the best experts on this subject based on the ideXlab platform.
-
CPM - Efficient Discovery of Proximity Patterns with Suffix Arrays
2001Co-Authors: Hiroki Arimura, Hiroki Asaka, Hiroshi Sakamoto, Setsuo ArikawaAbstract:We describe an efficient implementation of a Text Mining Algorithm for discovering a class of simple string patterns. With an index structure, called the virtual suffix tree, for pattern discovery built on the top of the suffix array, the resulting Algorithm is simple and fast in practice compared with the previous implementation with the suffix tree.
-
Efficient Discovery of Proximity Patterns with Suffix Arrays (Extended Abstract)
Combinatorial Pattern Matching, 2001Co-Authors: Hiroki Arimura, Hiroki Asaka, Hiroshi Sakamoto, Setsuo ArikawaAbstract:We describe an efficient implementation of a Text Mining Algorithm for discovering a class of simple string patterns. With an index structure, called the virtual suffix tree, for pattern discovery built on the top of the suffix array, the resulting Algorithm is simple and fast in practice compared with the previous implementation with the suffix tree.
-
Efficient discovery of proximity patterns with suffix arrays
Lecture Notes in Computer Science, 2001Co-Authors: Hiroki Arimura, Hiroki Asaka, Hiroshi Sakamoto, Setsuo ArikawaAbstract:We describe an efficient implementation of a Text Mining Algorithm for discovering a class of simple string patterns. With an index structure, called the virtual suffix tree, for pattern discovery built on the top of the suffix array, the resulting Algorithm is simple and fast in practice compared with the previous implementation with the suffix tree.
Frances B. Maguire - One of the best experts on this subject based on the ideXlab platform.
-
A Text-Mining approach to obtain detailed treatment information from free-Text fields in population-based cancer registries: A study of non-small cell lung cancer in California
PloS one, 2019Co-Authors: Frances B. Maguire, Cyllene R. Morris, Arti Parikh-patel, Rosemary D. Cress, Theresa H. M. Keegan, Patrick S. Lin, Kenneth W. KizerAbstract:Author(s): Maguire, Frances B; Morris, Cyllene R; Parikh-Patel, Arti; Cress, Rosemary D; Keegan, Theresa H M; Li, Chin-Shang; Lin, Patrick S; Kizer, Kenneth W | Abstract: Population-based cancer registries have treatment information for all patients making them an excellent resource for population-level monitoring. However, specific treatment details, such as drug names, are contained in a free-Text format that is difficult to process and summarize. We assessed the accuracy and efficiency of a Text-Mining Algorithm to identify systemic treatments for lung cancer from free-Text fields in the California Cancer Registry.The Algorithm used Perl regular expressions in SAS 9.4 to search for treatments in 24,845 free-Text records associated with 17,310 patients in California diagnosed with stage IV non-small cell lung cancer between 2012 and 2014. Our Algorithm categorized treatments into six groups that align with National Comprehensive Cancer Network guidelines. We compared results to a manual review (gold standard) of the same records.Percent agreement ranged from 91.1% to 99.4%. Ranges for other measures were 0.71-0.92 (Kappa), 74.3%-97.3% (sensitivity), 92.4%-99.8% (specificity), 60.4%-96.4% (positive predictive value), and 92.9%-99.9% (negative predictive value). The Text-Mining Algorithm used one-sixth of the time required for manual review.SAS-based Text Mining of free-Text data can accurately detect systemic treatments administered to patients and save considerable time compared to manual review, maximizing the utility of the extant information in population-based cancer registries for comparative effectiveness research.
-
Agreement of treatment between the SAS Text Mining Algorithm and manual review among stage IV non-small cell lung cancer patients (n = 17, 310), 2012–2014, California.
2019Co-Authors: Frances B. Maguire, Cyllene R. Morris, Arti Parikh-patel, Rosemary D. Cress, Theresa H. M. Keegan, Patrick S. Lin, Kenneth W. KizerAbstract:Agreement of treatment between the SAS Text Mining Algorithm and manual review among stage IV non-small cell lung cancer patients (n = 17, 310), 2012–2014, California.