The Experts below are selected from a list of 9615 Experts worldwide ranked by ideXlab platform
Ali Selamat - One of the best experts on this subject based on the ideXlab platform.
-
Arabic Script web page language identifications using decision tree neural networks
Pattern Recognition, 2011Co-Authors: Ali SelamatAbstract:In this paper, we propose a hybrid approach of Arabic Scripts web page language identification based on decision tree and ARTMAP approaches. We use the decision tree approach to find the general identities of a web document, be it an Arabic Script-based or a non-Arabic-based. Then, we use the selected representations of identified pages from the decision tree approach as an input to the ARTMAP neural network for further verification of the diversity of languages detected by the algorithm. From our initial experiments, we found that, although the decision tree approach may achieve a higher accuracy than ARTMAP, the former may not be as reliable as the ARTMAP approach if the language used is extended to other types of Arabic Script web documents in different languages (e.g., Urdu, Arabic, Persian, etc.). Therefore, we propose this hybrid decision tree-ARTMAP approach in order to improve the performance of the Arabic Script language identification on web documents in a variety of languages. The result shows that the proposed approach has outperformed both decision tree and the default ARTMAP approaches.
-
Arabic Script web page language identification using hybrid knn method
International Journal of Computational Intelligence and Applications, 2009Co-Authors: Ali Selamat, Iimam Much Ibnu SubrotoAbstract:In this paper, we proposed hybrid-KNN methods on the Arabic Script web page language identification. One of the crucial tasks in the text-based language identification that utilizes the same Script is how to produce reliable features and how to deal with the huge number of languages in the world. Specifically, it has involved the issue of feature representation, feature selection, identification performance, retrieval performance, and noise tolerance performance. Therefore, there are a number of methods that have been evaluated in this work; k-nearest neighbor (KNN), support vector machine (SVM), backpropagation neural networks (BPNN), hybrid KNN-SVM, and KNN-BPNN, in order to justify the capability of the state-of-the-art methods. KNN is prominent in data clustering or data filtering, SVM and BPNN are well known in supervised classification, and we have proposed hybrid-KNN for noise removal on web page language identification. We have used the standard measurements which are accuracy, precision, recall and F1 measurements to evaluate the effectiveness of the proposed hybrid-KNN. From the experiment, we have observed that BPNN is able to produce precise identification if the data set given is clean. However, when increasing the level of noise in the training data, KNN-SVM performs better than KNN-BPNN against the misclassification data, even on the level of 50% noise. Therefore, it is proven that KNN-SVM produce promising identification performance, in which KNN is able to reduce the noise in the data set and SVM is reliable in the language identification.
-
Improved Letter Weighting Feature Selection on Arabic Script Language Identification
2009 First Asian Conference on Intelligent Information and Database Systems, 2009Co-Authors: Choon-ching Ng, Ali SelamatAbstract:Language identification is the process identifying predefined language in a document automatically; we focused on the Web documents in this paper. Initially, we have applied the letter frequency as features combine with neural networks in Arabic Script language identification. However, reliability of selected letters of the features is a major issue to be overcome. Therefore, we propose an improved letter weighting feature selection in order to enhance the effectiveness of language identification. It is based on the concept letter frequency document frequency. From the experiments, we have found that the improved letter weighting feature selection achieve the highest accuracy 99.75% on Arabic Script language identification.
-
ACIIDS - Improved Letter Weighting Feature Selection on Arabic Script Language Identification
2009 First Asian Conference on Intelligent Information and Database Systems, 2009Co-Authors: Choon-ching Ng, Ali SelamatAbstract:Language identification is the process identifying predefined language in a document automatically; we focused on the web documents in this paper. Initially, we have applied the letter frequency as features combine with neural networks in Arabic Script language identification. However, reliability of selected letters of the features is a major issue to be overcome. Therefore, we propose an improved letter weighting feature selection in order to enhance the effectiveness of language identification. It is based on the concept letter frequency document frequency. From the experiments, we have found that the improved letter weighting feature selection achieve the highest accuracy 99.75% on Arabic Script language identification.
-
Arabic Script language identification using letter frequency neural networks
International Journal of Web Information Systems, 2008Co-Authors: Ali SelamatAbstract:Purpose – With the rapid emergence and explosion of the internet and the trend of globalization, a tremendous number of textual documents written in different languages are electronically accessible online from the world wide web. Efficiently and effectively managing these documents written in different languages is important to organizations and individuals. Therefore, the purpose of this paper is to propose letter frequency neural networks to enhance the performance of language identification.Design/methodology/approach – Initially, the paper analyzes the feasibility of using a windowing algorithm in order to find the best method in selecting the features of Arabic Script documents language identification using backpropagation neural networks. Previously, it had been found that the sliding window and non‐sliding window algorithm used as feature selection methods in the experiments did not yield a good result. Therefore, this paper proposes, a language identification of Arabic Script documents based on l...
William F Clocksin - One of the best experts on this subject based on the ideXlab platform.
-
structural features of cursive Arabic Script
British Machine Vision Conference, 1999Co-Authors: Mohammad S Khorsheed, William F ClocksinAbstract:We present a technique for extracting structural features from cursive Arabic Script. After preprocessing, the skeleton of the binary word image is decomposed into a number of segments in a certain order. Each segment is transformed into a feature vector. The target features are the curvature of the segment, its length relative to other segment lengths of the same word, the position of the segment relative to the centroid of the skeleton, and detailed deScription of curved segments. The result of this method is used to train the Hidden Markov Model to perform the recognition.
-
BMVC - Structural Features of Cursive Arabic Script.
Procedings of the British Machine Vision Conference 1999, 1999Co-Authors: Mohammad S Khorsheed, William F ClocksinAbstract:We present a technique for extracting structural features from cursive Arabic Script. After preprocessing, the skeleton of the binary word image is decomposed into a number of segments in a certain order. Each segment is transformed into a feature vector. The target features are the curvature of the segment, its length relative to other segment lengths of the same word, the position of the segment relative to the centroid of the skeleton, and detailed deScription of curved segments. The result of this method is used to train the Hidden Markov Model to perform the recognition.
Raid Saabne - One of the best experts on this subject based on the ideXlab platform.
-
real time segmentation of on line handwritten Arabic Script
International Conference on Frontiers in Handwriting Recognition, 2014Co-Authors: George Kour, Raid SaabneAbstract:Abstract —Real-time performance is necessary inapplications involving on-line handwriting recognition.However, conventional approaches usually wait until the entirecurve is traced out before starting the analysis, inevitablycausing delays in the recognition process. In regards to theArabic Script, the postponed analysis may be attributed to thecursive and unconstrained nature of the Arabic writing system,in both printed and handwritten forms. Nevertheless, thispaper proposes a real-time recognition-based segmentationtechnique of on-line Arabic Script. It demonstrate thefeasibility of carrying out the most time consuming tasks,required for the segmentation process, during the course ofwriting. The system has been designed and tested using theADAB Database, and promising results were obtained. Keywords -Arabic Script segmentation; handwriting recogni-tion; on-line text segmentation; I. I NTRODUCTION Handwriting remains the most commonly used meanof communication and recording of information in thedaily life, therefore, a growing interest in the handwritingcharacter recognition field has emerged in recent years.Handwriting recognition can be categorized into two mainareas: off-line and on-line. In the off-line case, a digitalimage containing text is fed to the computer, and the systemattempts to convert the spatial representation of the lettersinto digital symbols [1]. In contrast, the process of on-linehandwriting recognition is done on a digital representation ofthe text written on a special digitizer, tablet or smart-phonedevice, where sensors pick up the pen-tip movements.Research in this field has established two main ap-proaches; the analytic approach, which involves segmen-tation and classification of each part of the text [2], [3],[4], and the holistic approach, which considers the globalproperties of the written text and recognizes the input wordshape as a whole [5], [6]. While having many advantages, theholistic approach requires the classifier to be trained over theentire dictionary, which is impractical for large dictionaries(containing more than 20,000 words) [7].The cursiveness of the Arabic Script, prima facie, requiresdelaying the launch of the recognition process until thecompletion of the word scribing. However, in this paper, wequestion the necessity of this requirement by demonstratingthe feasibility of approximating the position of the
-
ICFHR - Real-Time Segmentation of On-Line Handwritten Arabic Script
2014 14th International Conference on Frontiers in Handwriting Recognition, 2014Co-Authors: George Kour, Raid SaabneAbstract:Abstract —Real-time performance is necessary inapplications involving on-line handwriting recognition.However, conventional approaches usually wait until the entirecurve is traced out before starting the analysis, inevitablycausing delays in the recognition process. In regards to theArabic Script, the postponed analysis may be attributed to thecursive and unconstrained nature of the Arabic writing system,in both printed and handwritten forms. Nevertheless, thispaper proposes a real-time recognition-based segmentationtechnique of on-line Arabic Script. It demonstrate thefeasibility of carrying out the most time consuming tasks,required for the segmentation process, during the course ofwriting. The system has been designed and tested using theADAB Database, and promising results were obtained. Keywords -Arabic Script segmentation; handwriting recogni-tion; on-line text segmentation; I. I NTRODUCTION Handwriting remains the most commonly used meanof communication and recording of information in thedaily life, therefore, a growing interest in the handwritingcharacter recognition field has emerged in recent years.Handwriting recognition can be categorized into two mainareas: off-line and on-line. In the off-line case, a digitalimage containing text is fed to the computer, and the systemattempts to convert the spatial representation of the lettersinto digital symbols [1]. In contrast, the process of on-linehandwriting recognition is done on a digital representation ofthe text written on a special digitizer, tablet or smart-phonedevice, where sensors pick up the pen-tip movements.Research in this field has established two main ap-proaches; the analytic approach, which involves segmen-tation and classification of each part of the text [2], [3],[4], and the holistic approach, which considers the globalproperties of the written text and recognizes the input wordshape as a whole [5], [6]. While having many advantages, theholistic approach requires the classifier to be trained over theentire dictionary, which is impractical for large dictionaries(containing more than 20,000 words) [7].The cursiveness of the Arabic Script, prima facie, requiresdelaying the launch of the recognition process until thecompletion of the word scribing. However, in this paper, wequestion the necessity of this requirement by demonstratingthe feasibility of approximating the position of the
Mohammad S Khorsheed - One of the best experts on this subject based on the ideXlab platform.
-
structural features of cursive Arabic Script
British Machine Vision Conference, 1999Co-Authors: Mohammad S Khorsheed, William F ClocksinAbstract:We present a technique for extracting structural features from cursive Arabic Script. After preprocessing, the skeleton of the binary word image is decomposed into a number of segments in a certain order. Each segment is transformed into a feature vector. The target features are the curvature of the segment, its length relative to other segment lengths of the same word, the position of the segment relative to the centroid of the skeleton, and detailed deScription of curved segments. The result of this method is used to train the Hidden Markov Model to perform the recognition.
-
BMVC - Structural Features of Cursive Arabic Script.
Procedings of the British Machine Vision Conference 1999, 1999Co-Authors: Mohammad S Khorsheed, William F ClocksinAbstract:We present a technique for extracting structural features from cursive Arabic Script. After preprocessing, the skeleton of the binary word image is decomposed into a number of segments in a certain order. Each segment is transformed into a feature vector. The target features are the curvature of the segment, its length relative to other segment lengths of the same word, the position of the segment relative to the centroid of the skeleton, and detailed deScription of curved segments. The result of this method is used to train the Hidden Markov Model to perform the recognition.
C. V. Jawahar - One of the best experts on this subject based on the ideXlab platform.
-
Unconstrained Scene Text and Video Text Recognition for Arabic Script
arXiv: Computer Vision and Pattern Recognition, 2017Co-Authors: Mohit Jain, Minesh Mathew, C. V. JawaharAbstract:Building robust recognizers for Arabic has always been challenging. We demonstrate the effectiveness of an end-to-end trainable CNN-RNN hybrid architecture in recognizing Arabic text in videos and natural scenes. We outperform previous state-of-the-art on two publicly available video text datasets - ALIF and ACTIV. For the scene text recognition task, we introduce a new Arabic scene text dataset and establish baseline results. For Scripts like Arabic, a major challenge in developing robust recognizers is the lack of large quantity of annotated data. We overcome this by synthesising millions of Arabic text images from a large vocabulary of Arabic words and phrases. Our implementation is built on top of the model introduced here [37] which is proven quite effective for English scene text recognition. The model follows a segmentation-free, sequence to sequence tranScription approach. The network transcribes a sequence of convolutional features from the input image to a sequence of target labels. This does away with the need for segmenting input image into constituent characters/glyphs, which is often difficult for Arabic Script. Further, the ability of RNNs to model contextual dependencies yields superior recognition results.
-
ASAR - Unconstrained scene text and video text recognition for Arabic Script
2017 1st International Workshop on Arabic Script Analysis and Recognition (ASAR), 2017Co-Authors: Mohit Jain, Minesh Mathew, C. V. JawaharAbstract:Building robust recognizers for Arabic has always been challenging. We demonstrate the effectiveness of an end-to-end trainable CNN-RNN hybrid architecture in recognizing Arabic text in videos and natural scenes. We outperform previous state-of-the-art on two publicly available video text datasets — ALIF and ACTIV. For the scene text recognition task, we introduce a new Arabic scene text dataset and establish baseline results. For Scripts like Arabic, a major challenge in developing robust recognizers is the lack of large quantity of annotated data. We overcome this by synthesizing millions of Arabic text images from a large vocabulary of Arabic words and phrases. Our implementation is built on top of the model introduced here [37] which is proven quite effective for English scene text recognition. The model follows a segmentation-free, sequence to sequence tranScription approach. The network transcribes a sequence of convolutional features from the input image to a sequence of target labels. This does away with the need for segmenting input image into constituent characters/glyphs, which is often difficult for Arabic Script. Further, the ability of RNNs to model contextual dependencies yields superior recognition results.