The Experts below are selected from a list of 315 Experts worldwide ranked by ideXlab platform
Thanaruk Theeramunkong - One of the best experts on this subject based on the ideXlab platform.
-
AINA Workshops - Improving Thai Academic Web Page Classification Using Inverse Class Frequency and Web Link Information
22nd International Conference on Advanced Information Networking and Applications - Workshops (aina workshops 2008), 2008Co-Authors: Verayuth Lertnattee, Thanaruk TheeramunkongAbstract:Automatic text Classification for Web collection is a non- trivial task. Since Thai academic Web pages usually present technical articles. They may have many technical terms both in Thai and English. This paper presents two approaches towards the problem of a large number of unique terms in a Web page: 1) term weighting schemes and 2) schemes using Web link information. We propose an approach using inverse Class Frequency instead of inverse document Frequency in centroid-based text categorization. Web link information provides information for users to follow to another part or page. It adds useful unique terms for Classification. The experimental results show that inverse Class Frequency is useful on a set of Thai academic Web documents, which is categorized by sources (sites) of information. It should be applied on both prototype and query vectors. Moreover, Web link information expresses its usefulness when inverse Class Frequency is also applied.
-
Analysis of inverse Class Frequency in centroid-based text Classification
IEEE International Symposium on Communications and Information Technology 2004. ISCIT 2004., 1Co-Authors: Verayuth Lertnattee, Thanaruk TheeramunkongAbstract:Most previous works on text categorization applied term occurrence Frequency and inverse document Frequency for representing importance of terms. This work presents an analysis of inverse Class Frequency in centroid-based text categorization. There are two aims of this paper. The first one is to find appropriate functions of inverse Class Frequency. The other is to find the key factors for using inverse Class Frequency. The experimental results show that the key factors, which improve Classification accuracy, are the numbers of few-Class terms and most-Class terms. When large numbers of few-Class terms and most-Class terms are obtained, the logarithmic function of inverse Class Frequency is the most effective when it is combined with term Frequency. The square root of inverse Class Frequency incorporated into TFIDF, works well in the case when data sets include a small number of few-Class terms and most-Class terms. To increase the numbers of these effective terms, some methods are involved i.e. using higher gram models, small number of Classes and large number of training sets.
-
Improving Thai educational Web page Classification using inverse Class Frequency
IEEE International Symposium on Communications and Information Technology 2005. ISCIT 2005., 1Co-Authors: Verayuth Lertnattee, Thanaruk TheeramunkongAbstract:Automatic text Classification for a Web collection is a challenge task, especially in the case that the language is not English, such as Thai. However, most of Thai educational Web pages usually include English terms due to their technical aspect. Lots of technical terms and typing errors both in Thai and in English are found in Web sites of universities. Most previous works on text categorization applied term Frequency and inverse document Frequency for representing importance of terms. In this paper, we use inverse Class Frequency instead of inverse document Frequency in centroid-based text categorization because it works well on a collection with a large number of unique terms. The experimental results show that inverse Class Frequency is useful, especially when it is applied on both prototype and query vectors.
Robert Lew - One of the best experts on this subject based on the ideXlab platform.
-
The Role of Syntactic Class, Frequency, and Word Order in Looking up English Multi-Word Expressions
Lexikos, 2012Co-Authors: Robert LewAbstract:Multi-word lexical units, such as compounds and idioms, are often problematic for lexicographers. Dictionaries are traditionally organized around single orthographic words, and so the question arises of where to place such complex lexical units. The user-friendly answer would be to include them primarily under the word which users are most likely to look up. But how do we know which words are likely to be looked up? The present study addresses this question by examining the roles of part of speech, word Frequency, and word position in guiding the decisions of Polish learners of English as to which component word of a multi-word expression to look up in the dictionary. The degree of word Frequency is found to be the strongest predictor, with less frequent words having a significantly greater chance of being selected for consultation. Then there is an independent part of speech-related preference for nouns, with adjectives being second, followed by verbs in third place. Words belonging to the remaining syntactic categories (adverbs, prepositions, conjunctions, determiners, and pronouns) are hardly looked up at all. However, word placement within the multi-word expression does not seem to matter much. This study has implications for dictionary makers in considering how to list multi-word-expressions.
-
The Role of Syntactic Class, Frequency, and Word Order in Looking up English
2012Co-Authors: Multi-word Expressions, Robert LewAbstract:Multi-word lexical units, such as compounds and idioms, are often problematic for lexicographers. Dictionaries are traditionally organized around single orthographic words, and so the question arises of where to place such complex lexical units. The user-friendly answer would be to include them primarily under the word which users are most likely to look up. But how do we know which words are likely to be looked up? The present study addresses this question by examining the roles of part of speech, word Frequency, and word position in guiding the decisions of Polish learners of English as to which component word of a multi-word expression to look up in the dictionary. The degree of word Frequency is found to be the strongest predictor, with less fre- quent words having a significantly greater chance of being selected for consultation. Then there is an independent part of speech-related preference for nouns, with adjectives being second, followed by verbs in third place. Words belonging to the remaining syntactic categories (adverbs, preposi- tions, conjunctions, determiners, and pronouns) are hardly looked up at all. However, word placement within the multi-word expression does not seem to matter much. This study has impli- cations for dictionary makers in considering how to list multi-word-expressions.
Carlo Semenza - One of the best experts on this subject based on the ideXlab platform.
-
when two and too don t go together a selective phonological deficit sparing number words
Cortex, 2011Co-Authors: Giulia Bencini, Lucia Pozzan, Laura Bertella, Ileana Mori, Riccardo Pignatti, Francesca Ceriani, Carlo SemenzaAbstract:Abstract We report the case of an Italian speaker (GBC) with Classical Wernicke’s aphasia syndrome following a vascular lesion in the left posterior middle temporal region. GBC exhibited a selective phonological deficit in spoken language production (repetition and reading) which affected all word Classes irrespective of grammatical Class, Frequency, and length. GBC’s production of number words, in contrast, was error free. The specific pattern of phonological errors on non-number words allows us to attribute the locus of impairment at the level of phonological form retrieval of a correctly selected lexical entry. These data support the claim that number words are represented and processed differently from other word categories in language production.
Daniel Ocaña - One of the best experts on this subject based on the ideXlab platform.
-
Mangrove vegetation assessment in the Santiago River Mouth, Mexico, by means of supervised Classification using LandsatTM imagery
Forest Ecology and Management, 1998Co-Authors: Pedro Ramírez-garcía, Jorge López-blanco, Daniel OcañaAbstract:This paper presents a mangrove vegetation assessment from 1970 to 1993 of the Santiago River Mouth, Nayarit, West of Mexico. The aims of this work are to describe the plant composition and structure of mangrove in the study area, and to evaluate the deforestation level and its amplitude by means of a retrospective analysis of the cover and distribution area of mangrove species using a LandsatTM image, aerial photographs and oblique video. Mangrove of the study area is dominated by Laguncularia racemosa with the average importance value of 158.18 and 400 ha of plant cover, followed by Avicennia germinans, with an average importance value of 138.52 and 324 ha of plant cover. Mangrove showed seven height and five diametrical Classes that include the two dominant species. L. racemosa was the dominant species in six of the eight compass lines. The highest absolute frequencies for both dominant species were found in the second height Class Frequency, and the first diametric Class Frequency. Cover area and distribution of mangrove in the study area were mapped using a LandsatTM5 image (April 1993). A supervised Classification was applied using the maximum likelihood algorithm, considering ten initial Classes. This Classification was evaluated by obtaining a Classification error matrix and by assessing its accuracy. The mangrove vegetation area reported before, considering the same area for image analysis, resulted to be overestimated in 56% regarding the value obtained in our photointerpretation (1065 ha). From the latter mangrove area, the current cover is 724 ha, which represents a decrease of 32% in a 23-yr period.
Verayuth Lertnattee - One of the best experts on this subject based on the ideXlab platform.
-
AINA Workshops - Improving Thai Academic Web Page Classification Using Inverse Class Frequency and Web Link Information
22nd International Conference on Advanced Information Networking and Applications - Workshops (aina workshops 2008), 2008Co-Authors: Verayuth Lertnattee, Thanaruk TheeramunkongAbstract:Automatic text Classification for Web collection is a non- trivial task. Since Thai academic Web pages usually present technical articles. They may have many technical terms both in Thai and English. This paper presents two approaches towards the problem of a large number of unique terms in a Web page: 1) term weighting schemes and 2) schemes using Web link information. We propose an approach using inverse Class Frequency instead of inverse document Frequency in centroid-based text categorization. Web link information provides information for users to follow to another part or page. It adds useful unique terms for Classification. The experimental results show that inverse Class Frequency is useful on a set of Thai academic Web documents, which is categorized by sources (sites) of information. It should be applied on both prototype and query vectors. Moreover, Web link information expresses its usefulness when inverse Class Frequency is also applied.
-
Analysis of inverse Class Frequency in centroid-based text Classification
IEEE International Symposium on Communications and Information Technology 2004. ISCIT 2004., 1Co-Authors: Verayuth Lertnattee, Thanaruk TheeramunkongAbstract:Most previous works on text categorization applied term occurrence Frequency and inverse document Frequency for representing importance of terms. This work presents an analysis of inverse Class Frequency in centroid-based text categorization. There are two aims of this paper. The first one is to find appropriate functions of inverse Class Frequency. The other is to find the key factors for using inverse Class Frequency. The experimental results show that the key factors, which improve Classification accuracy, are the numbers of few-Class terms and most-Class terms. When large numbers of few-Class terms and most-Class terms are obtained, the logarithmic function of inverse Class Frequency is the most effective when it is combined with term Frequency. The square root of inverse Class Frequency incorporated into TFIDF, works well in the case when data sets include a small number of few-Class terms and most-Class terms. To increase the numbers of these effective terms, some methods are involved i.e. using higher gram models, small number of Classes and large number of training sets.
-
Improving Thai educational Web page Classification using inverse Class Frequency
IEEE International Symposium on Communications and Information Technology 2005. ISCIT 2005., 1Co-Authors: Verayuth Lertnattee, Thanaruk TheeramunkongAbstract:Automatic text Classification for a Web collection is a challenge task, especially in the case that the language is not English, such as Thai. However, most of Thai educational Web pages usually include English terms due to their technical aspect. Lots of technical terms and typing errors both in Thai and in English are found in Web sites of universities. Most previous works on text categorization applied term Frequency and inverse document Frequency for representing importance of terms. In this paper, we use inverse Class Frequency instead of inverse document Frequency in centroid-based text categorization because it works well on a collection with a large number of unique terms. The experimental results show that inverse Class Frequency is useful, especially when it is applied on both prototype and query vectors.