The Experts below are selected from a list of 37641 Experts worldwide ranked by ideXlab platform

Yan Zhang - One of the best experts on this subject based on the ideXlab platform.

  • moocon a framework for semi supervised Concept Extraction from mooc content
    Database Systems for Advanced Applications, 2017
    Co-Authors: Zhuoxuan Jiang, Yan Zhang
    Abstract:

    Recent years have witnessed the rapid development of Massive Open Online Courses (MOOCs). MOOC platforms not only offer a one-stop learning setting, but also aggregate a large number of courses with various kinds of textual content, e.g. video subtitles, quizzes and forum content. MOOCs are also regarded as a large-scale ‘knowledge base’ which covers various domains. However, all the contents generated by instructors and learners are unstructured. In order to process the data to be structured for further knowledge management and mining, the first step could be Concept Extraction. In this paper, we expect to utilize human knowledge through labeling data, and propose a framework for Concept Extraction based on machine learning methods. The framework is flexible to support semi-supervised learning, in order to alleviate human effort of labeling training data. Also course-agnostic features are designed for modeling cross-domain data. Experimental results demonstrate that only 10% labeled data can lead to acceptable performance, and the semi-supervised learning method is comparable to the supervised version under the consistent framework. We find the textual contents of various forms, i.e. subtitles, PPTs and questions, should be separately processed due to their formal difference. At last we evaluate a new task: identifying needs of Concept comprehension. Our framework can work well in doing identification on forum content while learning a model from subtitles.

  • DASFAA Workshops - MOOCon: A Framework for Semi-supervised Concept Extraction from MOOC Content
    Database Systems for Advanced Applications, 2017
    Co-Authors: Zhuoxuan Jiang, Yan Zhang
    Abstract:

    Recent years have witnessed the rapid development of Massive Open Online Courses (MOOCs). MOOC platforms not only offer a one-stop learning setting, but also aggregate a large number of courses with various kinds of textual content, e.g. video subtitles, quizzes and forum content. MOOCs are also regarded as a large-scale ‘knowledge base’ which covers various domains. However, all the contents generated by instructors and learners are unstructured. In order to process the data to be structured for further knowledge management and mining, the first step could be Concept Extraction. In this paper, we expect to utilize human knowledge through labeling data, and propose a framework for Concept Extraction based on machine learning methods. The framework is flexible to support semi-supervised learning, in order to alleviate human effort of labeling training data. Also course-agnostic features are designed for modeling cross-domain data. Experimental results demonstrate that only 10% labeled data can lead to acceptable performance, and the semi-supervised learning method is comparable to the supervised version under the consistent framework. We find the textual contents of various forms, i.e. subtitles, PPTs and questions, should be separately processed due to their formal difference. At last we evaluate a new task: identifying needs of Concept comprehension. Our framework can work well in doing identification on forum content while learning a model from subtitles.

Zhuoxuan Jiang - One of the best experts on this subject based on the ideXlab platform.

  • moocon a framework for semi supervised Concept Extraction from mooc content
    Database Systems for Advanced Applications, 2017
    Co-Authors: Zhuoxuan Jiang, Yan Zhang
    Abstract:

    Recent years have witnessed the rapid development of Massive Open Online Courses (MOOCs). MOOC platforms not only offer a one-stop learning setting, but also aggregate a large number of courses with various kinds of textual content, e.g. video subtitles, quizzes and forum content. MOOCs are also regarded as a large-scale ‘knowledge base’ which covers various domains. However, all the contents generated by instructors and learners are unstructured. In order to process the data to be structured for further knowledge management and mining, the first step could be Concept Extraction. In this paper, we expect to utilize human knowledge through labeling data, and propose a framework for Concept Extraction based on machine learning methods. The framework is flexible to support semi-supervised learning, in order to alleviate human effort of labeling training data. Also course-agnostic features are designed for modeling cross-domain data. Experimental results demonstrate that only 10% labeled data can lead to acceptable performance, and the semi-supervised learning method is comparable to the supervised version under the consistent framework. We find the textual contents of various forms, i.e. subtitles, PPTs and questions, should be separately processed due to their formal difference. At last we evaluate a new task: identifying needs of Concept comprehension. Our framework can work well in doing identification on forum content while learning a model from subtitles.

  • DASFAA Workshops - MOOCon: A Framework for Semi-supervised Concept Extraction from MOOC Content
    Database Systems for Advanced Applications, 2017
    Co-Authors: Zhuoxuan Jiang, Yan Zhang
    Abstract:

    Recent years have witnessed the rapid development of Massive Open Online Courses (MOOCs). MOOC platforms not only offer a one-stop learning setting, but also aggregate a large number of courses with various kinds of textual content, e.g. video subtitles, quizzes and forum content. MOOCs are also regarded as a large-scale ‘knowledge base’ which covers various domains. However, all the contents generated by instructors and learners are unstructured. In order to process the data to be structured for further knowledge management and mining, the first step could be Concept Extraction. In this paper, we expect to utilize human knowledge through labeling data, and propose a framework for Concept Extraction based on machine learning methods. The framework is flexible to support semi-supervised learning, in order to alleviate human effort of labeling training data. Also course-agnostic features are designed for modeling cross-domain data. Experimental results demonstrate that only 10% labeled data can lead to acceptable performance, and the semi-supervised learning method is comparable to the supervised version under the consistent framework. We find the textual contents of various forms, i.e. subtitles, PPTs and questions, should be separately processed due to their formal difference. At last we evaluate a new task: identifying needs of Concept comprehension. Our framework can work well in doing identification on forum content while learning a model from subtitles.

Matthew H Samore - One of the best experts on this subject based on the ideXlab platform.

  • ICIMTH - Nora: A Vocabulary Discovery Tool for Concept Extraction.
    Studies in health technology and informatics, 2015
    Co-Authors: Guy Divita, Matthew H Samore, Marjorie E. Carter, B.s. Begum Durgahee, Warren B. P. Pettey, Andrew Redd, Adi V. Gundlapalli
    Abstract:

    Coverage of terms in domain-specific terminologies and ontologies is often limited in controlled medical vocabularies. Creating and augmenting such terminologies is resource intensive. We developed Nora as an interactive tool to discover terminology from text corpora; the output can then be employed to refine and enhance natural language processing-based Concept Extraction tasks. Nora provides a visualization of chains of words foraged from word frequency indexes from a text corpus. Domain experts direct and curate chains that contain relevant terms, which are further curated to identify lexical variants. A test of Nora demonstrated an increase of a domain lexicon in homelessness and related psychosocial factors by 38%, yielding an additional 10% extracted Concepts.

  • sophia a expedient umls Concept Extraction annotator
    American Medical Informatics Association Annual Symposium, 2014
    Co-Authors: Guy Divita, Qing Treitler Zeng, Adiseshu Venkata Gundlapalli, Scott L Duvall, Jonathan R Nebeker, Matthew H Samore
    Abstract:

    An opportunity exists for meaningful Concept Extraction and indexing from large corpora of clinical notes in the Veterans Affairs (VA) electronic medical record. Currently available tools such as MetaMap, cTAKES and HITex do not scale up to address this big data need. Sophia, a rapid UMLS Concept Extraction annotator was developed to fulfill a mandate and address Extraction where high throughput is needed while preserving performance. We report on the development, testing and benchmarking of Sophia against MetaMap and cTAKEs. Sophia demonstrated improved performance on recall as compared to cTAKES and MetaMap (0.71 vs 0.66 and 0.38). The overall f-score was similar to cTAKES and an improvement over MetaMap (0.53 vs 0.57 and 0.43). With regard to speed of processing records, we noted Sophia to be several fold faster than cTAKES and the scaled-out MetaMap service. Sophia offers a viable alternative for high-throughput information Extraction tasks.

  • AMIA - Sophia: A Expedient UMLS Concept Extraction Annotator.
    AMIA ... Annual Symposium proceedings. AMIA Symposium, 2014
    Co-Authors: Guy Divita, Qing Treitler Zeng, Adiseshu Venkata Gundlapalli, Scott L Duvall, Jonathan R Nebeker, Matthew H Samore
    Abstract:

    An opportunity exists for meaningful Concept Extraction and indexing from large corpora of clinical notes in the Veterans Affairs (VA) electronic medical record. Currently available tools such as MetaMap, cTAKES and HITex do not scale up to address this big data need. Sophia, a rapid UMLS Concept Extraction annotator was developed to fulfill a mandate and address Extraction where high throughput is needed while preserving performance. We report on the development, testing and benchmarking of Sophia against MetaMap and cTAKEs. Sophia demonstrated improved performance on recall as compared to cTAKES and MetaMap (0.71 vs 0.66 and 0.38). The overall f-score was similar to cTAKES and an improvement over MetaMap (0.53 vs 0.57 and 0.43). With regard to speed of processing records, we noted Sophia to be several fold faster than cTAKES and the scaled-out MetaMap service. Sophia offers a viable alternative for high-throughput information Extraction tasks.

Walter Daelemans - One of the best experts on this subject based on the ideXlab platform.

  • Unsupervised Concept Extraction from clinical text through semantic composition
    Journal of biomedical informatics, 2019
    Co-Authors: Stéphan Tulkens, Simon Suster, Walter Daelemans
    Abstract:

    Concept Extraction is an important step in clinical natural language processing. Once extracted, the use of Concepts can improve the accuracy and generalization of downstream systems. We present a new unsupervised system for the Extraction of Concepts from clinical text. The system creates representations of Concepts from the Unified Medical Language System (UMLS®) by combining natural language descriptions of Concepts with word representations, and composing these into higher-order Concept vectors. These Concept vectors are then used to assign labels to candidate phrases which are extracted using a syntactic chunker. Our approach scores an exact F-score of.32 and an inexact F-score of.45 on the well-known I2b2-2010 challenge corpus, outperforming the only other unsupervised Concept Extraction method. As our approach relies only on word representations and a chunker, it is completely unsupervised. As such, it can be applied to languages and corpora for which we do not have prior annotations. All our code is open-source and can be found at www.github.com/clips/conch.

M. Naphade - One of the best experts on this subject based on the ideXlab platform.

  • ICME - Multi-Modal Video Concept Extraction Using Co-Training
    2005 IEEE International Conference on Multimedia and Expo, 1
    Co-Authors: Rong Yan, M. Naphade
    Abstract:

    For large scale automatic semantic video characterization, it is necessary to learn and model a large number of semantic Concepts. A major obstacle to this is the insufficiency of labeled training samples. Semi-supervised learning algorithms such as co-training may help by incorporating a large amount of unlabeled data, which allows the redundant information across views to improve the learning performance. Although co-training has been successfully applied in several domains, it has not been used to detect video Concepts before. In this paper, we extend co-training to the domain of video Concept detection and investigate different strategies of co-training as well as their effects to the detection accuracy. We demonstrate performance based on the guideline of the TRECVID '03 semantic Concept Extraction task