The Experts below are selected from a list of 6447 Experts worldwide ranked by ideXlab platform

Chenglin Liu - One of the best experts on this subject based on the ideXlab platform.

  • handwritten mathematical expression recognition via paired adversarial learning
    International Journal of Computer Vision, 2020
    Co-Authors: Fei Yin, Yanming Zhang, Xuyao Zhang, Chenglin Liu
    Abstract:

    Recognition of handwritten mathematical expressions (MEs) is an important problem that has wide applications in practice. Handwritten ME recognition is challenging due to the variety of writing styles and ME formats. As a result, Recognizers trained by optimizing the traditional supervision loss do not perform satisfactorily. To improve the robustness of the recognizer with respect to writing styles, in this work, we propose a novel paired adversarial learning method to learn semantic-invariant features. Specifically, our proposed model, named PAL-v2, consists of an attention-based recognizer and a discriminator. During training, handwritten MEs and their printed templates are fed into PAL-v2 simultaneously. The attention-based recognizer is trained to learn semantic-invariant features with the guide of the discriminator. Moreover, we adopt a convolutional decoder to alleviate the vanishing and exploding gradient problems of RNN-based decoder, and further, improve the coverage of decoding with a novel attention method. We conducted extensive experiments on the CROHME dataset to demonstrate the effectiveness of each part of the method and achieved state-of-the-art performance.

Chungkwong Chan - One of the best experts on this subject based on the ideXlab platform.

  • stroke extraction for offline handwritten mathematical expression recognition
    IEEE Access, 2020
    Co-Authors: Chungkwong Chan
    Abstract:

    Offline handwritten mathematical expression recognition is often considered much harder than its online counterpart due to the absence of temporal information. In order to take advantage of the more mature methods for online recognition and save resources, an oversegmentation approach is proposed to recover strokes from textual bitmap images automatically. The proposed algorithm first breaks down the skeleton of a binarized image into junctions and segments, then segments are merged to form strokes, finally stroke order is normalized by using recursive projection and topological sort. Good offline accuracy was obtained in combination with ordinary online Recognizers, which were not specially designed for extracted strokes. Given a ready-made state-of-the-art online handwritten mathematical expression recognizer, the proposed procedure correctly recognized 58.22%, 65.65%, and 65.22% of the offline formulas rendered from the datasets of the Competitions on Recognition of Online Handwritten Mathematical Expressions (CROHME) in 2014, 2016, and 2019 respectively. Furthermore, given a trainable online recognition system, retraining it with extracted strokes resulted in an offline recognizer with the same level of accuracy. On the other hand, the speed of the entire pipeline was fast enough to facilitate on-device recognition on mobile phones with limited resources. To conclude, stroke extraction provides an attractive way to build optical character recognition software.

Eng Siong Chng - One of the best experts on this subject based on the ideXlab platform.

  • a target oriented phonotactic front end for spoken language recognition
    IEEE Transactions on Audio Speech and Language Processing, 2009
    Co-Authors: Rong Tong, Eng Siong Chng
    Abstract:

    This paper presents a strategy to optimize the phonotactic front-end for spoken language recognition. This is achieved by selecting a subset of phones from an existing phone recognizer's phone inventory such that only the phones that best discriminate each of the target languages are selected. Each such phone subset will be used to construct a target-oriented phone tokenizer (TOPT). In this study, we examine different approaches to construct such phone tokenizers for the front-end of a parallel phone Recognizers followed by vector space modeling (PPR-VSM) system. We show that the target-oriented phone tokenizers derived from language-specific phone Recognizers are more effective than the original parallel phone Recognizers. Our experimental results also show that the target-oriented phone tokenizers derived from universal phone Recognizers achieve better performance than those derived from language-specific phone Recognizers. Using the proposed target-oriented phone tokenizers as the phonotactic front-end, the language recognition system performance is significantly improved without the need for additional training samples. We achieve an equal error rate (EER) of 1.27%, 1.42% and 2.73% on the NIST 1996, 2003 and 2007 LRE databases respectively for 30-s closed-set tests. This system is one of the subsystems in IIR's submission to NIST 2007 LRE.

  • target oriented phone tokenizers for spoken language recognition
    International Conference on Acoustics Speech and Signal Processing, 2008
    Co-Authors: Rong Tong, Eng Siong Chng
    Abstract:

    This paper presents a new strategy for designing the parallel phone Recognizers for spoken language recognition. Given a collection of parallel phone Recognizers, we select a subset of phones from each phone recognizer for each target language to construct a target-oriented phone tokenizer (TOPT). As a result, the collection of target-oriented phone tokenizers is more effective than the original parallel phone Recognizers. This approach improves system performance significantly without requesting for additional transcribed training samples. We validate the effectiveness of the proposed strategy within the framework of the parallel phone recognizer followed by vector space modeling backend, or PPR-VSM. We achieve equal-error-rate of 2.21% and 3.65% on the 2003 and 2005 NIST LRE databases, respectively, for 30-second trials.

Bartišiūtė Gintarė - One of the best experts on this subject based on the ideXlab platform.

  • Medical – pharmaceutical information system with recognition of Lithuanian voice commands
    2020
    Co-Authors: Rudžionis Vytautas, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents a Lithuanian voice recognition system of medical - pharmaceutical terms. The system consists of two separate speech recognition modules working in parallel. One recognizer is a proprietary CD-HMM Lithuanian speech recognizer. The second recognizer is a Spanish speech recognizer adapted to recognize Lithuanian voice commands. The outputs of both Recognizers are combined by the decision making block yielding the final decision. The decision making block was automatically derived by an induction algorithm that learns a set of symbolic rules. The investigations showed that both Recognizers produce uncorrelated outputs and could complement each other. The investigations also showed that Lithuanian speech recognizer achieves higher accuracy (over 96 percent in a speaker independent mode) but the use of the adapted foreign language recognizer allows increase this baseline accuracy even further (over 98 percent in a speaker independent mode for 1000 voice commands). The voice recognition system is in the process of being embedded into several medical information systems which will be used by healthcare practitionersInformatikos fakultetasKauno technologijos universitetasSistemų analizės katedraVilnius university Kaunas faculty, LithuaniaVytauto Didžiojo universiteta

  • Medical – pharmaceutical information system with recognition of Lithuanian voice commands
    'IOS Press', 2020
    Co-Authors: Rudžionis Vytautas, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents a Lithuanian voice recognition system of medical - pharmaceutical terms. The system consists of two separate speech recognition modules working in parallel. One recognizer is a proprietary CD-HMM Lithuanian speech recognizer. The second recognizer is a Spanish speech recognizer adapted to recognize Lithuanian voice commands. The outputs of both Recognizers are combined by the decision making block yielding the final decision. The decision making block was automatically derived by an induction algorithm that learns a set of symbolic rules. The investigations showed that both Recognizers produce uncorrelated outputs and could complement each other. The investigations also showed that Lithuanian speech recognizer achieves higher accuracy (over 96 percent in a speaker independent mode) but the use of the adapted foreign language recognizer allows increase this baseline accuracy even further (over 98 percent in a speaker independent mode for 1000 voice commands). The voice recognition system is in the process of being embedded into several medical information systems which will be used by healthcare practitionersKauno technologijos universitetasSistemų analizės katedraVilnius university Kaunas faculty, LithuaniaVytauto Didžiojo universiteta

  • Web services based hybrid recognizer of Lithuanian voice commands
    'Kaunas University of Technology (KTU)', 2020
    Co-Authors: Rudžionis, Vytautas Evaldas, Maskeliūnas Rytis, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents the recently developed medical-pharmaceutical informative system with voice user interface. This is the first computerized system oriented towards healthcare services and industry where Lithuanian voice commands are used as a primary mean for control. Another essential property of the developed system is its hybrid nature: two different Recognizers - an adapted commercial Spanish speech recognizer available from Microsoft and a locally developed HMM speech recognizer based on Lithuanian acoustic models – are operating in parallel. The recognition hypotheses produced by those Recognizers are joined together using logical rules obtained using decision rules induction algorithms such as Ripper. All these measures and approaches allowed achieve very high speaker independent voice commands recognition accuracy acceptable for the system implementation in practice. The best achieved recognition was 98.9 % for 1000 Lithuanian voice commands. The paper presents optimization issues related with the development of the systemKauno technologijos universitetasSistemų analizės katedraVilniaus universitetas, vytautas.rudzionis@khf.vu.ltVytauto Didžiojo universiteta

  • Web services based hybrid recognizer of Lithuanian voice commands
    2020
    Co-Authors: Rudžionis, Vytautas Evaldas, Maskeliūnas Rytis, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents the recently developed medical-pharmaceutical informative system with voice user interface. This is the first computerized system oriented towards healthcare services and industry where Lithuanian voice commands are used as a primary mean for control. Another essential property of the developed system is its hybrid nature: two different Recognizers - an adapted commercial Spanish speech recognizer available from Microsoft and a locally developed HMM speech recognizer based on Lithuanian acoustic models – are operating in parallel. The recognition hypotheses produced by those Recognizers are joined together using logical rules obtained using decision rules induction algorithms such as Ripper. All these measures and approaches allowed achieve very high speaker independent voice commands recognition accuracy acceptable for the system implementation in practice. The best achieved recognition was 98.9 % for 1000 Lithuanian voice commands. The paper presents optimization issues related with the development of the systemInformatikos fakultetasKauno technologijos universitetasSistemų analizės katedraVilniaus universitetas, vytautas.rudzionis@khf.vu.ltVytauto Didžiojo universiteta

Ratkevičius Kastytis - One of the best experts on this subject based on the ideXlab platform.

  • Comparative analysis of adapted foreign language and native Lithuanian speech Recognizers for voice user interface
    2020
    Co-Authors: Rudžionis, Vytautas Evaldas, Maskeliūnas Rytis, Rudžionis, Algimantas Aleksandras, Ratkevičius Kastytis
    Abstract:

    Paper presents research results obtained when building a speaker independent hybrid speech recognizer. This recognizer will be integrated as a phrase recognizer in a medical-pharmaceutical information system. The hybrid speech recognizer consists of two recognition components: an adapted commercial Microsoft Spanish speech recognizer and a locally developed hidden Markov models based recognizer implementing Lithuanian acoustic models. Efficiency of both recognition components was evaluated on multiple speaker independent speech recognition tasks. The average accuracy of Lithuanian recognizer was higher reaching 0.6% phrase error rate for user requests in medical-pharmaceutical domain. The adapted commercial Spanish speech recognizer showed the ability to improve the accuracy of Lithuanian recognizer in the worst recognition scenarios. These results proved the hypothesis formulated when proposing the basic idea of hybrid recognition approach: recognition errors from different Recognizers built using various techniques are not strongly correlated. This fact could be exploited for improved overall speech recognition accuracyInformatikos fakultetasKauno technologijos universitetasSistemų analizės katedraVilniaus universitetasVytauto Didžiojo universiteta

  • Medical – pharmaceutical information system with recognition of Lithuanian voice commands
    2020
    Co-Authors: Rudžionis Vytautas, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents a Lithuanian voice recognition system of medical - pharmaceutical terms. The system consists of two separate speech recognition modules working in parallel. One recognizer is a proprietary CD-HMM Lithuanian speech recognizer. The second recognizer is a Spanish speech recognizer adapted to recognize Lithuanian voice commands. The outputs of both Recognizers are combined by the decision making block yielding the final decision. The decision making block was automatically derived by an induction algorithm that learns a set of symbolic rules. The investigations showed that both Recognizers produce uncorrelated outputs and could complement each other. The investigations also showed that Lithuanian speech recognizer achieves higher accuracy (over 96 percent in a speaker independent mode) but the use of the adapted foreign language recognizer allows increase this baseline accuracy even further (over 98 percent in a speaker independent mode for 1000 voice commands). The voice recognition system is in the process of being embedded into several medical information systems which will be used by healthcare practitionersInformatikos fakultetasKauno technologijos universitetasSistemų analizės katedraVilnius university Kaunas faculty, LithuaniaVytauto Didžiojo universiteta

  • Medical – pharmaceutical information system with recognition of Lithuanian voice commands
    'IOS Press', 2020
    Co-Authors: Rudžionis Vytautas, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents a Lithuanian voice recognition system of medical - pharmaceutical terms. The system consists of two separate speech recognition modules working in parallel. One recognizer is a proprietary CD-HMM Lithuanian speech recognizer. The second recognizer is a Spanish speech recognizer adapted to recognize Lithuanian voice commands. The outputs of both Recognizers are combined by the decision making block yielding the final decision. The decision making block was automatically derived by an induction algorithm that learns a set of symbolic rules. The investigations showed that both Recognizers produce uncorrelated outputs and could complement each other. The investigations also showed that Lithuanian speech recognizer achieves higher accuracy (over 96 percent in a speaker independent mode) but the use of the adapted foreign language recognizer allows increase this baseline accuracy even further (over 98 percent in a speaker independent mode for 1000 voice commands). The voice recognition system is in the process of being embedded into several medical information systems which will be used by healthcare practitionersKauno technologijos universitetasSistemų analizės katedraVilnius university Kaunas faculty, LithuaniaVytauto Didžiojo universiteta

  • Comparative analysis of adapted foreign language and native Lithuanian speech Recognizers for voice user interface
    2020
    Co-Authors: Rudžionis, Vytautas Evaldas, Maskeliūnas Rytis, Rudžionis, Algimantas Aleksandras, Ratkevičius Kastytis
    Abstract:

    Paper presents research results obtained when building a speaker independent hybrid speech recognizer. This recognizer will be integrated as a phrase recognizer in a medical-pharmaceutical information system. The hybrid speech recognizer consists of two recognition components: an adapted commercial Microsoft Spanish speech recognizer and a locally developed hidden Markov models based recognizer implementing Lithuanian acoustic models. Efficiency of both recognition components was evaluated on multiple speaker independent speech recognition tasks. The average accuracy of Lithuanian recognizer was higher reaching 0.6% phrase error rate for user requests in medical-pharmaceutical domain. The adapted commercial Spanish speech recognizer showed the ability to improve the accuracy of Lithuanian recognizer in the worst recognition scenarios. These results proved the hypothesis formulated when proposing the basic idea of hybrid recognition approach: recognition errors from different Recognizers built using various techniques are not strongly correlated. This fact could be exploited for improved overall speech recognition accuracyKauno technologijos universitetasSistemų analizės katedraVilniaus universitetasVytauto Didžiojo universiteta

  • Web services based hybrid recognizer of Lithuanian voice commands
    'Kaunas University of Technology (KTU)', 2020
    Co-Authors: Rudžionis, Vytautas Evaldas, Maskeliūnas Rytis, Ratkevičius Kastytis, Raškinis Gailius, Rudžionis Algimantas, Bartišiūtė Gintarė
    Abstract:

    This paper presents the recently developed medical-pharmaceutical informative system with voice user interface. This is the first computerized system oriented towards healthcare services and industry where Lithuanian voice commands are used as a primary mean for control. Another essential property of the developed system is its hybrid nature: two different Recognizers - an adapted commercial Spanish speech recognizer available from Microsoft and a locally developed HMM speech recognizer based on Lithuanian acoustic models – are operating in parallel. The recognition hypotheses produced by those Recognizers are joined together using logical rules obtained using decision rules induction algorithms such as Ripper. All these measures and approaches allowed achieve very high speaker independent voice commands recognition accuracy acceptable for the system implementation in practice. The best achieved recognition was 98.9 % for 1000 Lithuanian voice commands. The paper presents optimization issues related with the development of the systemKauno technologijos universitetasSistemų analizės katedraVilniaus universitetas, vytautas.rudzionis@khf.vu.ltVytauto Didžiojo universiteta