The Experts below are selected from a list of 147 Experts worldwide ranked by ideXlab platform
Sofiane Medjkoune - One of the best experts on this subject based on the ideXlab platform.
-
ICFHR - Mathematical Symbol Hypothesis Recognition with Rejection Option
2014 14th International Conference on Frontiers in Handwriting Recognition, 2014Co-Authors: Frank D. Julca-aguilar, Christian Viard-gaudin, Harold Mouchere, Nina S.t. Hirata, Sofiane MedjkouneAbstract:In the context of handwritten Mathematical expressions recognition, a first step consist on grouping strokes (segmentation) to form Symbol hypotheses: groups of strokes that might represent a Symbol. Then, the Symbol recognition step needs to cope with the identification of wrong segmented Symbols (false hypotheses). However, previous works on Symbol recognition consider only correctly segmented Symbols. In this work, we focus on the problem of Mathematical Symbol recognition where false hypotheses need to be identified. We extract Symbol hypotheses from complete handwritten Mathematical expressions and train artificial neural networks to perform both Symbol classification of true hypotheses and rejection of false hypotheses. We propose a new shape context-based Symbol descriptor: fuzzy shape context. Evaluation is performed on a publicly available dataset that contains 101 Symbol classes. Results show that the fuzzy shape context version outperforms the original shape context. Best recognition and false acceptance rates were obtained using a combination of shape contexts and online features: 86% and 17.5% respectively. As false rejection rate, we obtained 8.6% using only online features.
-
Mathematical Symbol Hypothesis Recognition with Rejection Option
2014 14th International Conference on Frontiers in Handwriting Recognition, 2014Co-Authors: Frank Julcaaguilar, Harold Mouchere, Nina S.t. Hirata, Christian Viardgaudin, Sofiane MedjkouneAbstract:In the context of handwritten Mathematical expressions recognition, a first step consist on grouping strokes (segmentation) to form Symbol hypotheses: groups of strokes that might represent a Symbol. Then, the Symbol recognition step needs to cope with the identification of wrong segmented Symbols (false hypotheses). However, previous works on Symbol recognition consider only correctly segmented Symbols. In this work, we focus on the problem of Mathematical Symbol recognition where false hypotheses need to be identified. We extract Symbol hypotheses from complete handwritten Mathematical expressions and train artificial neural networks to perform both Symbol classification of true hypotheses and rejection of false hypotheses. We propose a new shape context-based Symbol descriptor: fuzzy shape context. Evaluation is performed on a publicly available dataset that contains 101 Symbol classes. Results show that the fuzzy shape context version outperforms the original shape context. Best recognition and false acceptance rates were obtained using a combination of shape contexts and online features: 86% and 17.5% respectively. As false rejection rate, we obtained 8.6% using only online features.
-
ICDAR - Handwritten and Audio Information Fusion for Mathematical Symbol Recognition
2011 International Conference on Document Analysis and Recognition, 2011Co-Authors: Sofiane Medjkoune, Harold Mouchere, Simon Petitrenaud, Christian Viard-gaudinAbstract:Considerable efforts are being done within the scientific community to make as easier as possible the way that the human being converses with its machine. Handwriting and speech are two common ways used to achieve this goal and are probably among those which attracted much interest. In Mathematical content recognition tasks, these two modalities are used with a certain success. This paper presents an architecture based on a speech handwriting data fusion for isolated Mathematical Symbol recognition. Different fusion methods are explored. The results are very encouraging since recognition rates are increased comparatively to mono modality approaches.
-
Handwritten and Audio Information Fusion for Mathematical Symbol Recognition
2011 International Conference on Document Analysis and Recognition, 2011Co-Authors: Sofiane Medjkoune, Harold Mouchere, Simon Petitrenaud, Christian Viard-gaudinAbstract:Considerable efforts are being done within the scientific community to make as easier as possible the way that the human being converses with its machine. Handwriting and speech are two common ways used to achieve this goal and are probably among those which attracted much interest. In Mathematical content recognition tasks, these two modalities are used with a certain success. This paper presents an architecture based on a speech handwriting data fusion for isolated Mathematical Symbol recognition. Different fusion methods are explored. The results are very encouraging since recognition rates are increased comparatively to mono modality approaches.
Giovanni Soda - One of the best experts on this subject based on the ideXlab platform.
-
ICDAR - Using Earth Mover's Distance in the Bag-of-Visual-Words Model for Mathematical Symbol Retrieval
2011 International Conference on Document Analysis and Recognition, 2011Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper, the Earth Mover's Distance (EMD) is used as a similarity measure in the Mathematical Symbol retrieval task. The approach is based on the Bag-of-Visual-Words model. In our case the features extracted from each Symbol are clustered by means of Self-Organizing Maps (SOM) and then occurrences of features in the clusters are accumulated in a vector of visual words. The comparison between the latter vectors is performed with the EMD which naturally allows to incorporate the topological organization of SOM clusters in the distance computation. The proposed approach is experimentally tested in a Mathematical Symbol retrieval task and compared with the cosine similarity and with some variants that have been recently proposed.
-
Using Earth Mover's Distance in the Bag-of-Visual-Words Model for Mathematical Symbol Retrieval
2011 International Conference on Document Analysis and Recognition, 2011Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper, the Earth Mover's Distance (EMD) is used as a similarity measure in the Mathematical Symbol retrieval task. The approach is based on the Bag-of-Visual-Words model. In our case the features extracted from each Symbol are clustered by means of Self-Organizing Maps (SOM) and then occurrences of features in the clusters are accumulated in a vector of visual words. The comparison between the latter vectors is performed with the EMD which naturally allows to incorporate the topological organization of SOM clusters in the distance computation. The proposed approach is experimentally tested in a Mathematical Symbol retrieval task and compared with the cosine similarity and with some variants that have been recently proposed.
-
Mathematical Symbol indexing for digital libraries
Italian Research Conference on Digital Library Management Systems, 2010Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper we describe our recent research for Mathematical Symbol indexing and its possible application in the Digital Library domain. The proposed approach represents Mathematical Symbols by means of Shape Contexts (SC) description. Indexed Symbols are represented with a vector space-based method, but peculiar to our approach is the use of Self Organizing Maps (SOM) to perform the clustering instead of the commonly used k-means algorithm. The retrieval performance are measured on a large collection of Mathematical Symbols gathered from the widely used INFTY database.
-
IRCDL - Mathematical Symbol Indexing for Digital Libraries
Communications in Computer and Information Science, 2010Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper we describe our recent research for Mathematical Symbol indexing and its possible application in the Digital Library domain. The proposed approach represents Mathematical Symbols by means of Shape Contexts (SC) description. Indexed Symbols are represented with a vector space-based method, but peculiar to our approach is the use of Self Organizing Maps (SOM) to perform the clustering instead of the commonly used k-means algorithm. The retrieval performance are measured on a large collection of Mathematical Symbols gathered from the widely used INFTY database.
-
Mathematical Symbol indexing
Congress of the Italian Association for Artificial Intelligence, 2009Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:This paper addresses the indexing and retrieval of Mathematical Symbols from digitized documents. The proposed approach exploits Shape Contexts (SC) to describe the shape of Mathematical Symbols. Indexed Symbols are represented with a vector space-based method that is grounded on SC clustering. We explore the use of the Self Organizing Map (SOM) to perform the clustering and we compare several approaches to compute the SCs. The retrieval performance are measured on a large collection of Mathematical Symbols gathered from the widely used INFTY database.
Simone Marinai - One of the best experts on this subject based on the ideXlab platform.
-
ICDAR - Using Earth Mover's Distance in the Bag-of-Visual-Words Model for Mathematical Symbol Retrieval
2011 International Conference on Document Analysis and Recognition, 2011Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper, the Earth Mover's Distance (EMD) is used as a similarity measure in the Mathematical Symbol retrieval task. The approach is based on the Bag-of-Visual-Words model. In our case the features extracted from each Symbol are clustered by means of Self-Organizing Maps (SOM) and then occurrences of features in the clusters are accumulated in a vector of visual words. The comparison between the latter vectors is performed with the EMD which naturally allows to incorporate the topological organization of SOM clusters in the distance computation. The proposed approach is experimentally tested in a Mathematical Symbol retrieval task and compared with the cosine similarity and with some variants that have been recently proposed.
-
Using Earth Mover's Distance in the Bag-of-Visual-Words Model for Mathematical Symbol Retrieval
2011 International Conference on Document Analysis and Recognition, 2011Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper, the Earth Mover's Distance (EMD) is used as a similarity measure in the Mathematical Symbol retrieval task. The approach is based on the Bag-of-Visual-Words model. In our case the features extracted from each Symbol are clustered by means of Self-Organizing Maps (SOM) and then occurrences of features in the clusters are accumulated in a vector of visual words. The comparison between the latter vectors is performed with the EMD which naturally allows to incorporate the topological organization of SOM clusters in the distance computation. The proposed approach is experimentally tested in a Mathematical Symbol retrieval task and compared with the cosine similarity and with some variants that have been recently proposed.
-
Mathematical Symbol indexing for digital libraries
Italian Research Conference on Digital Library Management Systems, 2010Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper we describe our recent research for Mathematical Symbol indexing and its possible application in the Digital Library domain. The proposed approach represents Mathematical Symbols by means of Shape Contexts (SC) description. Indexed Symbols are represented with a vector space-based method, but peculiar to our approach is the use of Self Organizing Maps (SOM) to perform the clustering instead of the commonly used k-means algorithm. The retrieval performance are measured on a large collection of Mathematical Symbols gathered from the widely used INFTY database.
-
IRCDL - Mathematical Symbol Indexing for Digital Libraries
Communications in Computer and Information Science, 2010Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:In this paper we describe our recent research for Mathematical Symbol indexing and its possible application in the Digital Library domain. The proposed approach represents Mathematical Symbols by means of Shape Contexts (SC) description. Indexed Symbols are represented with a vector space-based method, but peculiar to our approach is the use of Self Organizing Maps (SOM) to perform the clustering instead of the commonly used k-means algorithm. The retrieval performance are measured on a large collection of Mathematical Symbols gathered from the widely used INFTY database.
-
Mathematical Symbol indexing
Congress of the Italian Association for Artificial Intelligence, 2009Co-Authors: Simone Marinai, Beatrice Miotti, Giovanni SodaAbstract:This paper addresses the indexing and retrieval of Mathematical Symbols from digitized documents. The proposed approach exploits Shape Contexts (SC) to describe the shape of Mathematical Symbols. Indexed Symbols are represented with a vector space-based method that is grounded on SC clustering. We explore the use of the Self Organizing Map (SOM) to perform the clustering and we compare several approaches to compute the SCs. The retrieval performance are measured on a large collection of Mathematical Symbols gathered from the widely used INFTY database.
Masakazu Suzuki - One of the best experts on this subject based on the ideXlab platform.
-
Detecting Mathematical Expressions in Scientific Document Images Using a U-Net Trained on a Diverse Dataset
IEEE Access, 2019Co-Authors: Wataru Ohyama, Masakazu Suzuki, Seiichi UchidaAbstract:A detection method for Mathematical expressions in scientific document images is proposed. Inspired by the promising performance of U-Net, a convolutional network architecture originally proposed for the semantic segmentation of biomedical images, the proposed method uses image conversion by a U-Net framework. The proposed method does not use any information from Mathematical and linguistic grammar so that it can be a supplemental bypass in the conventional Mathematical optical character recognition (OCR) process pipeline. The evaluation experiments confirmed that (1) the performance of Mathematical Symbol and expression detection by the proposed method is superior to that of InftyReader, which is state-of-the-art software for Mathematical OCR; (2) the coverage of the training dataset to the variation of document style is important; and (3) retraining with small additional training samples will be effective to improve the performance. An additional contribution is the release of a dataset for benchmarking the OCR for scientific documents.
-
Mathematical Symbol recognition with support vector machines
Pattern Recognition Letters, 2008Co-Authors: Christopher Malon, Seiichi Uchida, Masakazu SuzukiAbstract:Single-character recognition of Mathematical Symbols poses challenges from its two-dimensional pattern, the variety of similar Symbols that must be recognized distinctly, the imbalance and paucity of training data available, and the impossibility of final verification through spell check. We investigate the use of support vector machines to improve the classification of InftyReader, a free system for the OCR of Mathematical documents. First, we compare the performance of SVM kernels and feature definitions on pairs of letters that InftyReader usually confuses. Second, we describe a successful approach to multi-class classification with SVM, utilizing the ranking of alternatives within InftyReader's confusion clusters. The inclusion of our technique in InftyReader reduces its misrecognition rate by 41%.
-
support vector machines for Mathematical Symbol recognition
Lecture Notes in Computer Science, 2006Co-Authors: Christopher Malon, Seiichi Uchida, Masakazu SuzukiAbstract:Mathematical formulas challenge an OCR system with a range of similar-looking characters whose bold, calligraphic, and italic varieties must be recognized distinctly, though the fonts to be used in an article are not known in advance. We describe the use of support vector machines (SVM) to learn and predict about 300 classes of styled characters and Symbols.
-
SSPR/SPR - Support vector machines for Mathematical Symbol recognition
Lecture Notes in Computer Science, 2006Co-Authors: Christopher Malon, Seiichi Uchida, Masakazu SuzukiAbstract:Mathematical formulas challenge an OCR system with a range of similar-looking characters whose bold, calligraphic, and italic varieties must be recognized distinctly, though the fonts to be used in an article are not known in advance. We describe the use of support vector machines (SVM) to learn and predict about 300 classes of styled characters and Symbols.
Seiichi Uchida - One of the best experts on this subject based on the ideXlab platform.
-
Detecting Mathematical Expressions in Scientific Document Images Using a U-Net Trained on a Diverse Dataset
IEEE Access, 2019Co-Authors: Wataru Ohyama, Masakazu Suzuki, Seiichi UchidaAbstract:A detection method for Mathematical expressions in scientific document images is proposed. Inspired by the promising performance of U-Net, a convolutional network architecture originally proposed for the semantic segmentation of biomedical images, the proposed method uses image conversion by a U-Net framework. The proposed method does not use any information from Mathematical and linguistic grammar so that it can be a supplemental bypass in the conventional Mathematical optical character recognition (OCR) process pipeline. The evaluation experiments confirmed that (1) the performance of Mathematical Symbol and expression detection by the proposed method is superior to that of InftyReader, which is state-of-the-art software for Mathematical OCR; (2) the coverage of the training dataset to the variation of document style is important; and (3) retraining with small additional training samples will be effective to improve the performance. An additional contribution is the release of a dataset for benchmarking the OCR for scientific documents.
-
Mathematical Symbol recognition with support vector machines
Pattern Recognition Letters, 2008Co-Authors: Christopher Malon, Seiichi Uchida, Masakazu SuzukiAbstract:Single-character recognition of Mathematical Symbols poses challenges from its two-dimensional pattern, the variety of similar Symbols that must be recognized distinctly, the imbalance and paucity of training data available, and the impossibility of final verification through spell check. We investigate the use of support vector machines to improve the classification of InftyReader, a free system for the OCR of Mathematical documents. First, we compare the performance of SVM kernels and feature definitions on pairs of letters that InftyReader usually confuses. Second, we describe a successful approach to multi-class classification with SVM, utilizing the ranking of alternatives within InftyReader's confusion clusters. The inclusion of our technique in InftyReader reduces its misrecognition rate by 41%.
-
support vector machines for Mathematical Symbol recognition
Lecture Notes in Computer Science, 2006Co-Authors: Christopher Malon, Seiichi Uchida, Masakazu SuzukiAbstract:Mathematical formulas challenge an OCR system with a range of similar-looking characters whose bold, calligraphic, and italic varieties must be recognized distinctly, though the fonts to be used in an article are not known in advance. We describe the use of support vector machines (SVM) to learn and predict about 300 classes of styled characters and Symbols.
-
SSPR/SPR - Support vector machines for Mathematical Symbol recognition
Lecture Notes in Computer Science, 2006Co-Authors: Christopher Malon, Seiichi Uchida, Masakazu SuzukiAbstract:Mathematical formulas challenge an OCR system with a range of similar-looking characters whose bold, calligraphic, and italic varieties must be recognized distinctly, though the fonts to be used in an article are not known in advance. We describe the use of support vector machines (SVM) to learn and predict about 300 classes of styled characters and Symbols.