The Experts below are selected from a list of 13995 Experts worldwide ranked by ideXlab platform
Hyuk-chul Kwon - One of the best experts on this subject based on the ideXlab platform.
-
Improved Gini-Index Algorithm to Correct Feature-Selection Bias in Text Classification
IEICE Transactions on Information and Systems, 2011Co-Authors: Heum Park, Hyuk-chul KwonAbstract:This paper presents an improved Gini-Index algorithm to correct feature-selection bias in text classification. Gini-Index has been used as a split measure for choosing the most appropriate splitting attribute in decision tree. Recently, an improved Gini-Index algorithm for feature selection, designed for text categorization and based on Gini-Index theory, was introduced, and it has proved to be better than the other methods. However, we found that the Gini-Index still shows a feature selection bias in text classification, specifically for unbalanced datasets having a huge number of features. The feature selection bias of the Gini-Index in feature selection is shown in three ways: 1) the Gini values of low-frequency features are low (on purity measure) overall, irrespective of the distribution of features among classes, 2) for high-frequency features, the Gini values are always relatively high and 3) for specific features belonging to large classes, the Gini values are relatively lower than those belonging to small classes. Therefore, to correct that bias and improve feature selection in text classification using Gini-Index, we propose an improved Gini-Index (I-GI) algorithm with three reformulated Gini-Index expressions. In the present study, we used global dimensionality reduction (DR) and local DR to measure the goodness of features in feature selections. In experimental results for the I-GI algorithm, we obtained unbiased feature values and eliminated many irrelevant general features while retaining many specific features. Furthermore, we could improve the overall classification performances when we used the local DR method. The total averages of the classification performance were increased by 19.4%, 15.9%, 3.3%, 2.8% and 2.9% (kNN) in Micro-F1, 14%, 9.8%, 9.2%, 3.5% and 4.3% (SVM) in Micro-F1, 20%, 16.9%, 2.8%, 3.6% and 3.1% (kNN) in Macro-F1, 16.3%, 14%, 7.1%, 4.4%, 6.3% (SVM) in Macro-F1, compared with tf*idf, χ2, Information Gain, Odds Ratio and the existing Gini-Index methods according to each classifier.
-
complete Gini Index text git feature selection algorithm for text classification
International Conference on Software Engineering, 2010Co-Authors: Heum Park, Soonho Kwon, Hyuk-chul KwonAbstract:The recently introduced Gini-Index Text (GIT) feature-selection algorithm for text classification, through incorporating an improved Gini Index for better feature-selection performance, has some drawbacks. Specifically, the algorithm, under real-world experimental conditions, concentrates feature values to one point and be inadequate for selecting representative features. As such, good representative features cannot be estimated, and neither, moreover, can good performance be achieved in unbalanced text classification. Therefore, we suggest a new complete GIT feature-selection algorithm for text classification. The new algorithm, according to experimental results, could obtain unbiased feature values, and could eliminate many irrelevant and redundant features from feature subsets while retaining many representative features. Furthermore, the new algorithm, compared with the original version, demonstrated a notably improved overall classification performance.
Tomson Ogwang - One of the best experts on this subject based on the ideXlab platform.
-
The Gini Index for a Quadratic Pen's Parade
2010Co-Authors: Tomson OgwangAbstract:In this paper, we provide alternative derivations of the Gini Index for a quadratic Pen's parade without imposing unnecessary restrictions on its parameters. It turns out that in sufficiently large samples the reference value of the Gini Index for a quadratic Pen's parade is 1/2 (as opposed to 1/3 in the case of a linear Pen's parade). Whether the Gini Index is equal to, less than or greater than the reference value in these cases depends on whether or not Pen's parade passes through, above or below the origin, respectively.
-
Additional properties of a linear pen's parade for individual data using the stochastic approach to the Gini Index
Economics Letters, 2007Co-Authors: Tomson OgwangAbstract:Abstract The stochastic approach to the Gini Index is used to derive new properties of a linear Pen's parade for individual data. With small sample adjustments, the reference value of the Gini Index is independent of the number of income-receiving units.
-
AN UPPER BOUND OF THE Gini Index IN THE ABSENCE OF MEAN INCOME INFORMATION
Review of Income and Wealth, 2006Co-Authors: Tomson OgwangAbstract:In this paper, an upper bound of the Gini Index, based on grouped data, is proposed assuming that there is no information on all the group mean incomes as well as the overall mean income but the limits of the income brackets are known. An important advantage of this proposal is that conventional formulas for the upper bound of the Gini Index could be applied directly by substituting the (unknown) mean income for each income bracket with the corresponding value that maximizes the grouping correction for that income bracket. The effects of varying the number and size of income brackets are investigated.
-
Minor Concentration Ratios for the Bounds of the Gini Index
Journal of Income Distribution, 2004Co-Authors: Tomson OgwangAbstract:The minor concentration ratio is used to supplement the Gini Index in income distribution studies. The appeal of the minor concentration ratio stems form the fact that it examines the relative position of the “poor”, an important focus group in the analysis of income distributions. In this note, minor concentration ratios associated with the lower and upper bounds of the Gini Index are derived based on the observed points of the Lorenz curve. When the two minor concentration ratios are computed using grouped data for the United States, they turn out to be fairly close.
Heum Park - One of the best experts on this subject based on the ideXlab platform.
-
Improved Gini-Index Algorithm to Correct Feature-Selection Bias in Text Classification
IEICE Transactions on Information and Systems, 2011Co-Authors: Heum Park, Hyuk-chul KwonAbstract:This paper presents an improved Gini-Index algorithm to correct feature-selection bias in text classification. Gini-Index has been used as a split measure for choosing the most appropriate splitting attribute in decision tree. Recently, an improved Gini-Index algorithm for feature selection, designed for text categorization and based on Gini-Index theory, was introduced, and it has proved to be better than the other methods. However, we found that the Gini-Index still shows a feature selection bias in text classification, specifically for unbalanced datasets having a huge number of features. The feature selection bias of the Gini-Index in feature selection is shown in three ways: 1) the Gini values of low-frequency features are low (on purity measure) overall, irrespective of the distribution of features among classes, 2) for high-frequency features, the Gini values are always relatively high and 3) for specific features belonging to large classes, the Gini values are relatively lower than those belonging to small classes. Therefore, to correct that bias and improve feature selection in text classification using Gini-Index, we propose an improved Gini-Index (I-GI) algorithm with three reformulated Gini-Index expressions. In the present study, we used global dimensionality reduction (DR) and local DR to measure the goodness of features in feature selections. In experimental results for the I-GI algorithm, we obtained unbiased feature values and eliminated many irrelevant general features while retaining many specific features. Furthermore, we could improve the overall classification performances when we used the local DR method. The total averages of the classification performance were increased by 19.4%, 15.9%, 3.3%, 2.8% and 2.9% (kNN) in Micro-F1, 14%, 9.8%, 9.2%, 3.5% and 4.3% (SVM) in Micro-F1, 20%, 16.9%, 2.8%, 3.6% and 3.1% (kNN) in Macro-F1, 16.3%, 14%, 7.1%, 4.4%, 6.3% (SVM) in Macro-F1, compared with tf*idf, χ2, Information Gain, Odds Ratio and the existing Gini-Index methods according to each classifier.
-
complete Gini Index text git feature selection algorithm for text classification
International Conference on Software Engineering, 2010Co-Authors: Heum Park, Soonho Kwon, Hyuk-chul KwonAbstract:The recently introduced Gini-Index Text (GIT) feature-selection algorithm for text classification, through incorporating an improved Gini Index for better feature-selection performance, has some drawbacks. Specifically, the algorithm, under real-world experimental conditions, concentrates feature values to one point and be inadequate for selecting representative features. As such, good representative features cannot be estimated, and neither, moreover, can good performance be achieved in unbalanced text classification. Therefore, we suggest a new complete GIT feature-selection algorithm for text classification. The new algorithm, according to experimental results, could obtain unbiased feature values, and could eliminate many irrelevant and redundant features from feature subsets while retaining many representative features. Furthermore, the new algorithm, compared with the original version, demonstrated a notably improved overall classification performance.
Ying-ju Chen - One of the best experts on this subject based on the ideXlab platform.
-
A new computational approach for estimation of the Gini Index based on grouped data
Computational Statistics, 2021Co-Authors: Tatjana Miljkovic, Ying-ju ChenAbstract:Many government agencies still rely on the grouped data as the main source of information for calculation of the Gini Index. Previous research showed that the Gini Index based on the grouped data suffers the first and second-order correction bias compared to the Gini Index computed based on the individual data. Since the accuracy of the estimated correction bias is subject to many underlying assumptions, we propose a new method and name it D-Gini, which reduces the bias in Gini coefficient based on grouped data. We investigate the performance of the D-Gini method on an open-ended tail interval of the income distribution. The results of our simulation study showed that our method is very effective in minimizing the first and second order-bias in the Gini Index and outperforms other methods previously used for the bias-correction of the Gini Index based on grouped data. Three data sets are used to illustrate the application of this method.
Amir Shoham - One of the best experts on this subject based on the ideXlab platform.
-
Practical modified Gini Index
Applied Economics Letters, 2013Co-Authors: Miki Malul, Daniel Shapira, Amir ShohamAbstract:The Gini Index is the most common method for estimating the level of income inequality in countries. In this article, we suggest a simple modification that takes into account the moderating effect of in-kind government benefits. Unlike other studies that use micro-level data that are rarely available for many countries or over a period of time, the proposed Modified Gini (MGini) Index could be calculated using just the regularly available data for each country. Such data include the original Gini coefficient, government consumption expenditures, Gross Domestic Product (GDP) and total tax revenue as a percentage of GDP. This modified version of the Gini Index allows us to calculate the level of inequality more precisely and make better comparisons between countries and over time.
-
Practical Modified Gini Index. ACES Working Papers, 2012
2012Co-Authors: Amir Shoham, Daniel Shapira, Miki MalulAbstract:The Gini Index is the most common method for estimating the level of income inequality in countries. In this paper we suggest a simple modification that takes into account the moderating effect of in-kind government benefits. Unlike other studies that use micro level data that is rarely available for many countries or over a period of time, the proposed modified Gini Index could be calculated using just the regularly available data for each country. Such data includes the original Gini coefficient, government consumption expenditures, GDP and total tax revenue as a percentage of GDP. This modified version of the Gini Index allows us to calculate the level of inequality more precisely, and make better comparisons between countries and over time.