The Experts below are selected from a list of 360 Experts worldwide ranked by ideXlab platform

Muhammad Imran Razzak - One of the best experts on this subject based on the ideXlab platform.

  • handwritten urdu Character Recognition using one dimensional blstm classifier
    Neural Computing and Applications, 2019
    Co-Authors: Saad Bin Ahmed, Salahuddin Swati, Muhammad Imran Razzak
    Abstract:

    The Recognition of cursive script is regarded as a subtle task in optical Character Recognition due to its varied representation. Every cursive script has different nature and associated challenges. As Urdu is one of cursive language that is derived from Arabic script, that is why it nearly shares the similar challenges and complexities but with more intensity. We can categorize Urdu and Arabic language on basis of its script they use. Urdu is mostly written in Nasta’liq style, whereas Arabic follows Naskh style of writing. This paper presents new and comprehensive Urdu handwritten offline database name Urdu-Nasta’liq handwritten dataset (UNHD). Currently, there is no standard and comprehensive Urdu handwritten dataset available publicly for researchers. The acquired dataset covers commonly used ligatures that were written by 500 writers with their natural handwriting on A4 size paper. UNHD is publically available and can be download form https://sites.google.com/site/researchonurdulanguage1/databases . We performed experiments using recurrent neural networks and reported a significant accuracy for handwritten Urdu Character Recognition.

  • lexicon reduction for urdu arabic script based Character Recognition a multilingual ocr
    Mehran University Research Journal of Engineering and Technology, 2016
    Co-Authors: Saeeda Naz, Arif Iqbal Umar, Muhammad Imran Razzak
    Abstract:

    Arabic script Character Recognition is challenging task due to complexity of the script and huge number of ligatures. We present a method for the development of multilingual Arabic script OCR (Optical Character Recognition) and lexicon reduction for Arabic Script and its derivative languages. The objective of the proposed method is to overcome the large dataset Urdu and similar scripts by using GCT (Ghost Character Theory) concept. Arabic and its sibling script languages share the similar Character dataset i.e. the Character set are difference in diacritic and writing styles like Naskh or Nasta'liq. Based on the proposed method, the lexicon for Arabic and Arabic script based languages can be minimized approximately up to 20 times. The proposed multilingual Arabic script OCR approach have been evaluated for online Arabic and its derivative language like Urdu using BPNN. The result showed that proposed method helps to not only the reduction of lexicon but also helps to develop the Multilanguage Character Recognition system for Arabic Script.

  • the optical Character Recognition of urdu like cursive scripts
    Pattern Recognition, 2014
    Co-Authors: Khizar Hayat, Muhammad Imran Razzak, Muhammad Waqas Anwar, Sajjad Ahmad Madani, Samee U Khan
    Abstract:

    We survey the optical Character Recognition (OCR) literature with reference to the Urdu-like cursive scripts. In particular, the Urdu, Pushto, and Sindhi languages are discussed, with the emphasis being on the Nasta'liq and Naskh scripts. Before detaining the OCR works, the peculiarities of the Urdu-like scripts are outlined, which are followed by the presentation of the available text image databases. For the sake of clarity, the various attempts are grouped into three parts, namely: (a) printed, (b) handwritten, and (c) online Character Recognition. Within each part, the works are analyzed par rapport a typical OCR pipeline with an emphasis on the preprocessing, segmentation, feature extraction, classification, and Recognition. HighlightsA literature review of the Nasta'liq and Naskh cursive script OCR.The peculiarities and challenges are described a priori.Printed, handwritten and online OCR efforts are being explored.Analyses based on the stages of a typical OCR pipeline.

  • hmm and fuzzy logic a hybrid approach for online urdu script based languages Character Recognition
    Knowledge Based Systems, 2010
    Co-Authors: Muhammad Imran Razzak, Abdel Belaïd, Fareeha Anwar, Syed Afaq Husain, Muhammad Sher
    Abstract:

    Urdu script-based languages' Character Recognition has some technical issues not existing in other languages and makes these languages more complicated. Segmentation-based Character Recognition approach for handwritten Urdu, both Nasta'liq and Nasakh script-based languages, incorporates number of overhead and very less accurate as compared to segmentation free. This paper presents a segmentation-free approach for Recognition of online Urdu handwritten script using hybrid classifier, HMM and fuzzy logic. Trained data set consisting of HMMs for each stroke is further classified into 62 sub-patterns based on the primary stroke shape at the beginning and end using fuzzy rule. Fuzzy linguistic variables based on language structure are used to model features and provide suitable result for large variation in handwritten strokes. Twenty-six time variant structural and statistical features are extracted for the base strokes. The fuzzy classification into sub-patterns increases the efficiency and decreases the computational complexity due to reduction in data set size. The hybrid HMM-fuzzy technique is efficient for large and complex data set. It provided 87.6% and 74.1% for Nasta'liq and Nasakh, respectively, on 1800 ligatures.

  • locally baseline detection for online arabic script based languages Character Recognition
    International Journal of Physical Sciences, 2010
    Co-Authors: Muhammad Imran Razzak, Muhammad Sher, S A Hussain
    Abstract:

    Baseline detection is one of the most important step in Character Recognition and has direct influence on Recognition result. Due to the complexity of the Urdu scripts based languages, handwritten Character Recognition is a very difficult task as compared to other languages. Baseline detection is one of the main issue and basic step of mostly preprocessing operations that is, normalization, skewness, secondary strokes segmentation and also in feature extraction. This paper presents a novel method of baseline detection for cursive handwritten Urdu script. The proposed approach is divided into three steps: diacritical marks segmentation, primary baseline estimation and local baseline estimation. The local baseline extraction is estimated using the features extracted from ending shape of the words. Due to structural difference between Nasta'liq and Naskh style, different rules are formed for baseline estimation. Key words:

Muhammad Sher - One of the best experts on this subject based on the ideXlab platform.

  • hmm and fuzzy logic a hybrid approach for online urdu script based languages Character Recognition
    Knowledge Based Systems, 2010
    Co-Authors: Muhammad Imran Razzak, Abdel Belaïd, Fareeha Anwar, Syed Afaq Husain, Muhammad Sher
    Abstract:

    Urdu script-based languages' Character Recognition has some technical issues not existing in other languages and makes these languages more complicated. Segmentation-based Character Recognition approach for handwritten Urdu, both Nasta'liq and Nasakh script-based languages, incorporates number of overhead and very less accurate as compared to segmentation free. This paper presents a segmentation-free approach for Recognition of online Urdu handwritten script using hybrid classifier, HMM and fuzzy logic. Trained data set consisting of HMMs for each stroke is further classified into 62 sub-patterns based on the primary stroke shape at the beginning and end using fuzzy rule. Fuzzy linguistic variables based on language structure are used to model features and provide suitable result for large variation in handwritten strokes. Twenty-six time variant structural and statistical features are extracted for the base strokes. The fuzzy classification into sub-patterns increases the efficiency and decreases the computational complexity due to reduction in data set size. The hybrid HMM-fuzzy technique is efficient for large and complex data set. It provided 87.6% and 74.1% for Nasta'liq and Nasakh, respectively, on 1800 ligatures.

  • locally baseline detection for online arabic script based languages Character Recognition
    International Journal of Physical Sciences, 2010
    Co-Authors: Muhammad Imran Razzak, Muhammad Sher, S A Hussain
    Abstract:

    Baseline detection is one of the most important step in Character Recognition and has direct influence on Recognition result. Due to the complexity of the Urdu scripts based languages, handwritten Character Recognition is a very difficult task as compared to other languages. Baseline detection is one of the main issue and basic step of mostly preprocessing operations that is, normalization, skewness, secondary strokes segmentation and also in feature extraction. This paper presents a novel method of baseline detection for cursive handwritten Urdu script. The proposed approach is divided into three steps: diacritical marks segmentation, primary baseline estimation and local baseline estimation. The local baseline extraction is estimated using the features extracted from ending shape of the words. Due to structural difference between Nasta'liq and Naskh style, different rules are formed for baseline estimation. Key words:

  • combining offline and online preprocessing for online urdu Character Recognition
    2009
    Co-Authors: Muhammad Imran Razzak, Syed Afaq Hussain, Muhammad Sher, Zeeshan Shafi Khan
    Abstract:

    Urdu online handwriting Recognition is a very challenging task due to its cursive nature. Pre-processing of the raw input strokes is crucial part for the success of Character Recognition system. The findings from online data are not enough for Recognition of Urdu due to the complexities of Urdu script. This paper describes the preprocessing steps for online Character Recognition. A novel technique is presented for preprocessing of Urdu online text in which both online and offline domain are used to remove the variations and to increase the efficiency of the Recognition system for online input. The proposed technique is also the necessary step towards Character Recognition, person identification, personality determination where input data is processed from all perspectives.

Anil K. Jain - One of the best experts on this subject based on the ideXlab platform.

  • template based online Character Recognition
    Pattern Recognition, 2001
    Co-Authors: Scott D Connell, Anil K. Jain
    Abstract:

    Abstract Handwriting is a common, natural form of communication for humans, and therefore it is useful to utilize this modality as a means of input to machines. One well-known method of classifying individual Characters or words is template matching. We demonstrate a template-based system for online Character Recognition where the number of representative templates is determined automatically. These templates can be viewed as representing different styles of writing a particular Character. The templates are then used as a reference for efficient classification using decision trees. Overall, our classifier achieves an 86.9% accuracy on a set of 17,928 alphanumeric Characters (36 classes; 10 digits and 26 lowercase letters) with a throughput of over 8 Characters per second on a 296 MHz Sun UltraSparc.

  • FEATURE EXTRACTION METHODS FOR Character Recognition--A SURVEY
    Pattern Recognition, 1996
    Co-Authors: O.d. Trier, Anil K. Jain, Torfinn Taxt
    Abstract:

    This paper presents an overview of feature extraction methods for off-line Recognition of segmented (isolated) Characters. Selection of a feature extraction method is probably the single most important factor in achieving high Recognition performance in Character Recognition systems. Different feature extraction methods are designed for different representations of the Characters, such as solid binary Characters, Character contours, skeletons (thinned Characters) or gray-level subimages of each individual Character. The feature extraction methods are discussed in terms of invariance properties, reconstructability and expected distortions and variability of the Characters. The problem of choosing the appropriate feature extraction method for a given application is also discussed. When a few promising feature extraction methods have been identified, they need to be evaluated experimentally to find the best method for the given application.

Samee U Khan - One of the best experts on this subject based on the ideXlab platform.

  • the optical Character Recognition of urdu like cursive scripts
    Pattern Recognition, 2014
    Co-Authors: Khizar Hayat, Muhammad Imran Razzak, Muhammad Waqas Anwar, Sajjad Ahmad Madani, Samee U Khan
    Abstract:

    We survey the optical Character Recognition (OCR) literature with reference to the Urdu-like cursive scripts. In particular, the Urdu, Pushto, and Sindhi languages are discussed, with the emphasis being on the Nasta'liq and Naskh scripts. Before detaining the OCR works, the peculiarities of the Urdu-like scripts are outlined, which are followed by the presentation of the available text image databases. For the sake of clarity, the various attempts are grouped into three parts, namely: (a) printed, (b) handwritten, and (c) online Character Recognition. Within each part, the works are analyzed par rapport a typical OCR pipeline with an emphasis on the preprocessing, segmentation, feature extraction, classification, and Recognition. HighlightsA literature review of the Nasta'liq and Naskh cursive script OCR.The peculiarities and challenges are described a priori.Printed, handwritten and online OCR efforts are being explored.Analyses based on the stages of a typical OCR pipeline.

Abdel Belaïd - One of the best experts on this subject based on the ideXlab platform.

  • hmm and fuzzy logic a hybrid approach for online urdu script based languages Character Recognition
    Knowledge Based Systems, 2010
    Co-Authors: Muhammad Imran Razzak, Abdel Belaïd, Fareeha Anwar, Syed Afaq Husain, Muhammad Sher
    Abstract:

    Urdu script-based languages' Character Recognition has some technical issues not existing in other languages and makes these languages more complicated. Segmentation-based Character Recognition approach for handwritten Urdu, both Nasta'liq and Nasakh script-based languages, incorporates number of overhead and very less accurate as compared to segmentation free. This paper presents a segmentation-free approach for Recognition of online Urdu handwritten script using hybrid classifier, HMM and fuzzy logic. Trained data set consisting of HMMs for each stroke is further classified into 62 sub-patterns based on the primary stroke shape at the beginning and end using fuzzy rule. Fuzzy linguistic variables based on language structure are used to model features and provide suitable result for large variation in handwritten strokes. Twenty-six time variant structural and statistical features are extracted for the base strokes. The fuzzy classification into sub-patterns increases the efficiency and decreases the computational complexity due to reduction in data set size. The hybrid HMM-fuzzy technique is efficient for large and complex data set. It provided 87.6% and 74.1% for Nasta'liq and Nasakh, respectively, on 1800 ligatures.

  • Effect of Ghost Character Theory on Arabic Script Based Languages Character Recognition
    2009
    Co-Authors: Mohamed Imran Razzak, Abdel Belaïd, Syed Afaq Hussain
    Abstract:

    Arabic script is used by more than 1/4th population of the world in the form of different languages like Arabic, Persian, Urdu, Sindhi, Pashto etc but each language have its own words meaning. The set of شhas 58 alphabets. Arabic script based languages Character Recognition is difficult task due to complexities involved in this script not exist in other script. The analysis of the Arabic script is very complicated due to its use of diacritical marks associated with each Character and written in many fonts and style. This script has gain very less intention by the researcher. This paper present a novel technique named Ghost Character Recognition Theory that will helps to develop a Multilanguage Character Recognition system for Arabic script based languages based on Ghost Character Theory. The main benefit of proposed approach is that it will works for all Arabic script based languages by doing effort for ghost Character (basic skeleton) and developing dictionary for every language. By handling all Arabic script based languages many issues will arise like Recognition rate as compared to system for specific languages, but in general it is not big issue for multilingual system and at the end we will get multilingual Character Recognition system.