The Experts below are selected from a list of 26961 Experts worldwide ranked by ideXlab platform

Sun Kim - One of the best experts on this subject based on the ideXlab platform.

  • TeamTat: a collaborative Text Annotation tool.
    Nucleic acids research, 2020
    Co-Authors: Rezarta Islamaj, Dongseop Kwon, Sun Kim
    Abstract:

    Manually annotated data is key to developing Text-mining and information-extraction algorithms. However, human Annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing Text Annotation tools may provide user-friendly interfaces to domain experts, limited support is available for figure display, project management, and multi-user team Annotation. In response, we developed TeamTat (https://www.teamtat.org), a web-based Annotation tool (local setup available), equipped to manage team Annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document Annotation, reflecting the entire production life cycle. Project managers can specify Annotation schema for entities and relations and select annotator(s) and distribute documents anonymously to prevent bias. Document input format can be plain Text, PDF or BioC (uploaded locally or automatically retrieved from PubMed/PMC), and output format is BioC with inline Annotations. TeamTat displays figures from the full Text for the annotator's convenience. Multiple users can work on the same document independently in their workspaces, and the team manager can track task completion. TeamTat provides corpus quality assessment via inter-annotator agreement statistics, and a user-friendly interface convenient for Annotation review and inter-annotator disagreement resolution to improve corpus quality.

  • TeamTat: a collaborative Text Annotation tool
    arXiv: Human-Computer Interaction, 2020
    Co-Authors: Rezarta Islamaj, Dongseop Kwon, Sun Kim
    Abstract:

    Manually annotated data is key to developing Text-mining and information-extraction algorithms. However, human Annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing Text Annotation tools may provide user-friendly interfaces to domain experts, limited support is available for image display, project management, and multi-user team Annotation. In response, we developed TeamTat (this http URL), a web-based Annotation tool (local setup available), equipped to manage team Annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document Annotation, reflecting the entire production life cycle. Project managers can specify Annotation schema for entities and relations and select annotator(s) and distribute documents anonymously to prevent bias. Document input format can be plain Text, PDF or BioC, (uploaded locally or automatically retrieved from PubMed or PMC), and output format is BioC with inline Annotations. TeamTat displays figures from the full Text for the annotators convenience. Multiple users can work on the same document independently in their workspaces, and the team manager can track task completion. TeamTat provides corpus-quality assessment via inter-annotator agreement statistics, and a user-friendly interface convenient for Annotation review and inter-annotator disagreement resolution to improve corpus quality.

Rezarta Islamaj - One of the best experts on this subject based on the ideXlab platform.

  • TeamTat: a collaborative Text Annotation tool.
    Nucleic acids research, 2020
    Co-Authors: Rezarta Islamaj, Dongseop Kwon, Sun Kim
    Abstract:

    Manually annotated data is key to developing Text-mining and information-extraction algorithms. However, human Annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing Text Annotation tools may provide user-friendly interfaces to domain experts, limited support is available for figure display, project management, and multi-user team Annotation. In response, we developed TeamTat (https://www.teamtat.org), a web-based Annotation tool (local setup available), equipped to manage team Annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document Annotation, reflecting the entire production life cycle. Project managers can specify Annotation schema for entities and relations and select annotator(s) and distribute documents anonymously to prevent bias. Document input format can be plain Text, PDF or BioC (uploaded locally or automatically retrieved from PubMed/PMC), and output format is BioC with inline Annotations. TeamTat displays figures from the full Text for the annotator's convenience. Multiple users can work on the same document independently in their workspaces, and the team manager can track task completion. TeamTat provides corpus quality assessment via inter-annotator agreement statistics, and a user-friendly interface convenient for Annotation review and inter-annotator disagreement resolution to improve corpus quality.

  • TeamTat: a collaborative Text Annotation tool
    arXiv: Human-Computer Interaction, 2020
    Co-Authors: Rezarta Islamaj, Dongseop Kwon, Sun Kim
    Abstract:

    Manually annotated data is key to developing Text-mining and information-extraction algorithms. However, human Annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing Text Annotation tools may provide user-friendly interfaces to domain experts, limited support is available for image display, project management, and multi-user team Annotation. In response, we developed TeamTat (this http URL), a web-based Annotation tool (local setup available), equipped to manage team Annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document Annotation, reflecting the entire production life cycle. Project managers can specify Annotation schema for entities and relations and select annotator(s) and distribute documents anonymously to prevent bias. Document input format can be plain Text, PDF or BioC, (uploaded locally or automatically retrieved from PubMed or PMC), and output format is BioC with inline Annotations. TeamTat displays figures from the full Text for the annotators convenience. Multiple users can work on the same document independently in their workspaces, and the team manager can track task completion. TeamTat provides corpus-quality assessment via inter-annotator agreement statistics, and a user-friendly interface convenient for Annotation review and inter-annotator disagreement resolution to improve corpus quality.

Jiang Bian - One of the best experts on this subject based on the ideXlab platform.

  • User-centered design of a web-based crowdsourcing-integrated semantic Text Annotation tool for building a mental health knowledge base
    Journal of biomedical informatics, 2020
    Co-Authors: Hansi Zhang, Jiang Bian
    Abstract:

    Abstract Background One in five U.S. adults lives with some kind of mental health condition and 4.6% of all U.S. adults have a serious mental illness. The Internet has become the first place for these people to seek online mental health information for help. However, online mental health information is not well-organized and often of low quality. There have been efforts in building evidence-based mental health knowledgebases curated with information manually extracted from the high-quality scientific literature. Manual extraction is inefficient. Crowdsourcing can potentially be a low-cost mechanism to collect labeled data from non-expert laypeople. However, there is not an existing Annotation tool integrated with popular crowdsourcing platforms to perform the information extraction tasks. In our previous work, we prototyped a Semantic Text Annotation Tool (STAT) to address this gap. Objective We aimed to refine the STAT prototype (1) to improve its usability and (2) to enhance the crowdsourcing workflow efficiency to facilitate the construction of evidence-based mental health knowledgebase, following a user-centered design (UCD) approach. Methods Following UCD principles, we conducted four design iterations to improve the initial STAT prototype. In the first two iterations, usability testing focus groups were conducted internally with 8 participants recruited from a convenient sample, and the usability was evaluated with a modified System Usability Scale (SUS). In the following two iterations, usability testing was conducted externally using the Amazon Mechanical Turk (MTurk) platform. In each iteration, we summarized the usability testing results through thematic analysis, identified usability issues, and conducted a heuristic evaluation to map identified usability issues to Jakob Nielsen’s usability heuristics. We collected suggested improvements in the usability testing sessions and enhanced STAT accordingly in the next UCD iteration. After four UCD iterations, we conducted a case study of the system on MTurk using mental health related scientific literature. We compared the performance of crowdsourcing workers with two expert annotators from two aspects: efficiency and quality. Results The SUS score increased from 70.3 ± 12.5 to 81.1 ± 9.8 after the two internal UCD iterations as we improved STAT’s functionality based on the suggested improvements. We then evaluated STAT externally through MTurk in the following two iterations. The SUS score decreased to 55.7 ± 20.1 in the third iteration, probably because of the complexity of the tasks. After further simplification of STAT and the Annotation tasks with an improved Annotation guideline, the SUS score increased to 73.8 ± 13.8 in the fourth iteration of UCD. In the evaluation case study, on average, the workers spent 125.5 ± 69.2 s on the onboarding tutorial and the crowdsourcing workers spent significantly less time on the Annotation tasks compared to the two experts. In terms of Annotation quality, the workers’ Annotation results achieved average F1-scores ranged from 0.62 to 0.84 for the different sentences. Conclusions We successfully developed a web-based semantic Text Annotation tool, STAT, to facilitate the curation of semantic web knowledgebases through four UCD iterations. The lessons learned from the UCD process could serve as a guide to further enhance STAT and the development and design of other crowdsourcing-based semantic Text Annotation tasks. Our study also showed that a well-organized, informative Annotation guideline is as important as the Annotation tool itself. Further, we learned that a crowdsourcing task should consist of multiple simple microtasks rather than a complicated task.

  • Development of a Web-Based Crowdsourcing-Integrated Semantic Text Annotation Tool to Assist in Building a Mental Health Knowledge Base: User-Centered Design Approach (Preprint)
    2020
    Co-Authors: Hansi Zhang, Jiang Bian
    Abstract:

    BACKGROUND One in five U.S. adults lives with some kind of mental health condition and 4.6% of all U.S. adults have a serious mental illness in 2018. The Internet has become the first place for these people to seek online mental health information for help. However, online mental health information is not well-organized and often of low quality. There have been efforts in building evidence-based mental health knowledgebases curated with information manually extracted from the high-quality scientific literature. Manual extraction is inefficient. Crowdsourcing can potentially be a low-cost mechanism to collect labeled data from non-expert laypeople. However, there is not an existing Annotation tool integrated with popular crowdsourcing platforms to perform the information extraction tasks. In our previous work, we prototyped a Semantic Text Annotation Tool (STAT) to address this gap. OBJECTIVE We aimed to refine the STAT prototype (1) to improve its usability and (2) to enhance the crowdsourcing workflow efficiency to facilitate the construction of evidence-based mental health knowledgebase, following a user-centered design (UCD) process. METHODS Following UCD principles, we conducted four design iterations to improve the initial STAT prototype. In the first two iterations, usability testing focus groups were conducted internally with 8 participants recruited from a convenient sample, and the usability was evaluated with a modified System Usability Scale (SUS). In the following two iterations, usability testing was conducted externally using the Amazon Mechanical Turk (MTurk) platform. In each iteration, we summarized the usability testing results through thematic analysis, identified usability issues, and conducted a heuristic evaluation to map identified usability issues to Jakob Nielsen’s usability heuristics. We collected suggested improvements in each of the usability testing sessions and enhanced STAT accordingly in the next UCD iteration. After four UCD iterations, we conducted a case study of the system on MTurk using mental health related scientific literature. We compared the performance of crowdsourcing workers with two expert annotators from two aspects: efficiency and quality. RESULTS At the end of two initial internal UCD iterations, the SUS score increased from 70.3 ± 12.5 to 81.1 ± 9.8 after we improved STAT following the suggested improvements. We then evaluated STAT externally through MTurk in the following two iterations. The SUS score decreased to 55.7 ± 20.1 in the third iteration, probably because of the complexity of the tasks. After further simplification of STAT and the Annotation tasks with an improved Annotation guideline, the SUS score increased to 73.8 ± 13.8 in the fourth iteration of UCD. In the evaluation case study, on average, the workers spent 125.5 ± 69.2 seconds on the onboarding tutorial and the crowdsourcing workers spent significantly less time on the Annotation tasks compared to the two experts. In terms of Annotation quality, the workers’ Annotation results achieved average F1-scores ranged from 0.62 to 0.84 for the different sentences. CONCLUSIONS We successfully developed a web-based semantic Text Annotation tool, STAT, to facilitate the curation of semantic web knowledgebases through four UCD iterations. The lessons learned from the UCD process could serve as a guide to further enhance STAT and the development and design of other crowdsourcing-based semantic Text Annotation tasks. Our study also showed that a well-organized, informative Annotation guideline is as important as the Annotation tool itself. Further, we learned that a crowdsourcing task should consist of multiple simple microtasks rather than a complicated task.

  • ICHI - STAT: A Web-based Semantic Text Annotation Tool to Assist Building Mental Health Knowledge Base
    IEEE International Conference on Healthcare Informatics. IEEE International Conference on Healthcare Informatics, 2019
    Co-Authors: Hansi Zhang, Xi Yang, Yi Guo, Jiang Bian
    Abstract:

    Mental health problems are serious among American adults and many of them are turning to the Internet for help. However, online mental health information is not well-organized and in low quality. We are building a mental health knowledge base (MHKB) with evidence-based information extracted from scientific literature manually, but lacking efficiency. We envision to leverage collective wisdoms through crowdsourcing to speed up the curation of MHKB. In order to integrate with crowdsourcing platforms, we designed and prototyped a web-based Annotation tool, STAT (Semantic Text Annotation Tool), with real-time Annotation recommendation and Annotation quality analysis, to facilitate management of laypeople annotators recruited through crowdsourcing to complete the necessary Annotation tasks.

Dongseop Kwon - One of the best experts on this subject based on the ideXlab platform.

  • TeamTat: a collaborative Text Annotation tool.
    Nucleic acids research, 2020
    Co-Authors: Rezarta Islamaj, Dongseop Kwon, Sun Kim
    Abstract:

    Manually annotated data is key to developing Text-mining and information-extraction algorithms. However, human Annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing Text Annotation tools may provide user-friendly interfaces to domain experts, limited support is available for figure display, project management, and multi-user team Annotation. In response, we developed TeamTat (https://www.teamtat.org), a web-based Annotation tool (local setup available), equipped to manage team Annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document Annotation, reflecting the entire production life cycle. Project managers can specify Annotation schema for entities and relations and select annotator(s) and distribute documents anonymously to prevent bias. Document input format can be plain Text, PDF or BioC (uploaded locally or automatically retrieved from PubMed/PMC), and output format is BioC with inline Annotations. TeamTat displays figures from the full Text for the annotator's convenience. Multiple users can work on the same document independently in their workspaces, and the team manager can track task completion. TeamTat provides corpus quality assessment via inter-annotator agreement statistics, and a user-friendly interface convenient for Annotation review and inter-annotator disagreement resolution to improve corpus quality.

  • TeamTat: a collaborative Text Annotation tool
    arXiv: Human-Computer Interaction, 2020
    Co-Authors: Rezarta Islamaj, Dongseop Kwon, Sun Kim
    Abstract:

    Manually annotated data is key to developing Text-mining and information-extraction algorithms. However, human Annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing Text Annotation tools may provide user-friendly interfaces to domain experts, limited support is available for image display, project management, and multi-user team Annotation. In response, we developed TeamTat (this http URL), a web-based Annotation tool (local setup available), equipped to manage team Annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document Annotation, reflecting the entire production life cycle. Project managers can specify Annotation schema for entities and relations and select annotator(s) and distribute documents anonymously to prevent bias. Document input format can be plain Text, PDF or BioC, (uploaded locally or automatically retrieved from PubMed or PMC), and output format is BioC with inline Annotations. TeamTat displays figures from the full Text for the annotators convenience. Multiple users can work on the same document independently in their workspaces, and the team manager can track task completion. TeamTat provides corpus-quality assessment via inter-annotator agreement statistics, and a user-friendly interface convenient for Annotation review and inter-annotator disagreement resolution to improve corpus quality.

Hansi Zhang - One of the best experts on this subject based on the ideXlab platform.

  • User-centered design of a web-based crowdsourcing-integrated semantic Text Annotation tool for building a mental health knowledge base
    Journal of biomedical informatics, 2020
    Co-Authors: Hansi Zhang, Jiang Bian
    Abstract:

    Abstract Background One in five U.S. adults lives with some kind of mental health condition and 4.6% of all U.S. adults have a serious mental illness. The Internet has become the first place for these people to seek online mental health information for help. However, online mental health information is not well-organized and often of low quality. There have been efforts in building evidence-based mental health knowledgebases curated with information manually extracted from the high-quality scientific literature. Manual extraction is inefficient. Crowdsourcing can potentially be a low-cost mechanism to collect labeled data from non-expert laypeople. However, there is not an existing Annotation tool integrated with popular crowdsourcing platforms to perform the information extraction tasks. In our previous work, we prototyped a Semantic Text Annotation Tool (STAT) to address this gap. Objective We aimed to refine the STAT prototype (1) to improve its usability and (2) to enhance the crowdsourcing workflow efficiency to facilitate the construction of evidence-based mental health knowledgebase, following a user-centered design (UCD) approach. Methods Following UCD principles, we conducted four design iterations to improve the initial STAT prototype. In the first two iterations, usability testing focus groups were conducted internally with 8 participants recruited from a convenient sample, and the usability was evaluated with a modified System Usability Scale (SUS). In the following two iterations, usability testing was conducted externally using the Amazon Mechanical Turk (MTurk) platform. In each iteration, we summarized the usability testing results through thematic analysis, identified usability issues, and conducted a heuristic evaluation to map identified usability issues to Jakob Nielsen’s usability heuristics. We collected suggested improvements in the usability testing sessions and enhanced STAT accordingly in the next UCD iteration. After four UCD iterations, we conducted a case study of the system on MTurk using mental health related scientific literature. We compared the performance of crowdsourcing workers with two expert annotators from two aspects: efficiency and quality. Results The SUS score increased from 70.3 ± 12.5 to 81.1 ± 9.8 after the two internal UCD iterations as we improved STAT’s functionality based on the suggested improvements. We then evaluated STAT externally through MTurk in the following two iterations. The SUS score decreased to 55.7 ± 20.1 in the third iteration, probably because of the complexity of the tasks. After further simplification of STAT and the Annotation tasks with an improved Annotation guideline, the SUS score increased to 73.8 ± 13.8 in the fourth iteration of UCD. In the evaluation case study, on average, the workers spent 125.5 ± 69.2 s on the onboarding tutorial and the crowdsourcing workers spent significantly less time on the Annotation tasks compared to the two experts. In terms of Annotation quality, the workers’ Annotation results achieved average F1-scores ranged from 0.62 to 0.84 for the different sentences. Conclusions We successfully developed a web-based semantic Text Annotation tool, STAT, to facilitate the curation of semantic web knowledgebases through four UCD iterations. The lessons learned from the UCD process could serve as a guide to further enhance STAT and the development and design of other crowdsourcing-based semantic Text Annotation tasks. Our study also showed that a well-organized, informative Annotation guideline is as important as the Annotation tool itself. Further, we learned that a crowdsourcing task should consist of multiple simple microtasks rather than a complicated task.

  • Development of a Web-Based Crowdsourcing-Integrated Semantic Text Annotation Tool to Assist in Building a Mental Health Knowledge Base: User-Centered Design Approach (Preprint)
    2020
    Co-Authors: Hansi Zhang, Jiang Bian
    Abstract:

    BACKGROUND One in five U.S. adults lives with some kind of mental health condition and 4.6% of all U.S. adults have a serious mental illness in 2018. The Internet has become the first place for these people to seek online mental health information for help. However, online mental health information is not well-organized and often of low quality. There have been efforts in building evidence-based mental health knowledgebases curated with information manually extracted from the high-quality scientific literature. Manual extraction is inefficient. Crowdsourcing can potentially be a low-cost mechanism to collect labeled data from non-expert laypeople. However, there is not an existing Annotation tool integrated with popular crowdsourcing platforms to perform the information extraction tasks. In our previous work, we prototyped a Semantic Text Annotation Tool (STAT) to address this gap. OBJECTIVE We aimed to refine the STAT prototype (1) to improve its usability and (2) to enhance the crowdsourcing workflow efficiency to facilitate the construction of evidence-based mental health knowledgebase, following a user-centered design (UCD) process. METHODS Following UCD principles, we conducted four design iterations to improve the initial STAT prototype. In the first two iterations, usability testing focus groups were conducted internally with 8 participants recruited from a convenient sample, and the usability was evaluated with a modified System Usability Scale (SUS). In the following two iterations, usability testing was conducted externally using the Amazon Mechanical Turk (MTurk) platform. In each iteration, we summarized the usability testing results through thematic analysis, identified usability issues, and conducted a heuristic evaluation to map identified usability issues to Jakob Nielsen’s usability heuristics. We collected suggested improvements in each of the usability testing sessions and enhanced STAT accordingly in the next UCD iteration. After four UCD iterations, we conducted a case study of the system on MTurk using mental health related scientific literature. We compared the performance of crowdsourcing workers with two expert annotators from two aspects: efficiency and quality. RESULTS At the end of two initial internal UCD iterations, the SUS score increased from 70.3 ± 12.5 to 81.1 ± 9.8 after we improved STAT following the suggested improvements. We then evaluated STAT externally through MTurk in the following two iterations. The SUS score decreased to 55.7 ± 20.1 in the third iteration, probably because of the complexity of the tasks. After further simplification of STAT and the Annotation tasks with an improved Annotation guideline, the SUS score increased to 73.8 ± 13.8 in the fourth iteration of UCD. In the evaluation case study, on average, the workers spent 125.5 ± 69.2 seconds on the onboarding tutorial and the crowdsourcing workers spent significantly less time on the Annotation tasks compared to the two experts. In terms of Annotation quality, the workers’ Annotation results achieved average F1-scores ranged from 0.62 to 0.84 for the different sentences. CONCLUSIONS We successfully developed a web-based semantic Text Annotation tool, STAT, to facilitate the curation of semantic web knowledgebases through four UCD iterations. The lessons learned from the UCD process could serve as a guide to further enhance STAT and the development and design of other crowdsourcing-based semantic Text Annotation tasks. Our study also showed that a well-organized, informative Annotation guideline is as important as the Annotation tool itself. Further, we learned that a crowdsourcing task should consist of multiple simple microtasks rather than a complicated task.

  • ICHI - STAT: A Web-based Semantic Text Annotation Tool to Assist Building Mental Health Knowledge Base
    IEEE International Conference on Healthcare Informatics. IEEE International Conference on Healthcare Informatics, 2019
    Co-Authors: Hansi Zhang, Xi Yang, Yi Guo, Jiang Bian
    Abstract:

    Mental health problems are serious among American adults and many of them are turning to the Internet for help. However, online mental health information is not well-organized and in low quality. We are building a mental health knowledge base (MHKB) with evidence-based information extracted from scientific literature manually, but lacking efficiency. We envision to leverage collective wisdoms through crowdsourcing to speed up the curation of MHKB. In order to integrate with crowdsourcing platforms, we designed and prototyped a web-based Annotation tool, STAT (Semantic Text Annotation Tool), with real-time Annotation recommendation and Annotation quality analysis, to facilitate management of laypeople annotators recruited through crowdsourcing to complete the necessary Annotation tasks.