The Experts below are selected from a list of 4956 Experts worldwide ranked by ideXlab platform

Jungsun Kim - One of the best experts on this subject based on the ideXlab platform.

  • Author name disambiguation using a graph model with Node Splitting and merging based on bibliographic information
    Scientometrics, 2014
    Co-Authors: D Shin, T. Kim, J Choi, Jungsun Kim
    Abstract:

    Author ambiguity mainly arises when several different authors express\ntheir names in the same way, generally known as the namesake problem,\nand also when the name of an author is expressed in many different ways,\nreferred to as the heteronymous name problem. These author ambiguity\nproblems have long been an obstacle to efficient information retrieval\nin digital libraries, causing incorrect identification of authors and\nimpeding correct classification of their publications. It is a\nnontrivial task to distinguish those authors, especially when there is\nvery limited information about them. In this paper, we propose a graph\nbased approach to author name disambiguation, where a graph model is\nconstructed using the co-author relations, and author ambiguity is\nresolved by graph operations such as vertex (or Node) Splitting and\nmerging based on the co-authorship. In our framework, called a Graph\nFramework for Author Disambiguation (GFAD), the namesake problem is\nsolved by Splitting an author vertex involved in multiple cycles of\ncoauthorship, and the heteronymous name problem is handled by merging\nmultiple author vertices having similar names if those vertices are\nconnected to a common vertex. Experiments were carried out with the real\nDBLP and Arnetminer collections and the performance of GFAD is compared\nwith three representative unsupervised author name disambiguation\nsystems. We confirm that GFAD shows better overall performance from the\nperspective of representative evaluation metrics. An additional\ncontribution is that we released the refined DBLP collection to the\npublic to facilitate organizing a performance benchmark for future\nsystems on author disambiguation.

Andreas Ziegler - One of the best experts on this subject based on the ideXlab platform.

  • On the Use of Harrell's C for Node Splitting in Random Survival Forests
    arXiv: Machine Learning, 2015
    Co-Authors: Matthias Schmid, Marvin N. Wright, Andreas Ziegler
    Abstract:

    Random forests are one of the most successful methods for statistical learning and prediction. Here we consider random survival forests (RSF), which are an extension of the original random forest method to right-censored outcome variables. RSF use the log-rank split criterion to form an ensemble of survival trees, the prediction accuracy of the ensemble estimate is subsequently evaluated by the concordance index for survival data ("Harrell's C"). Conceptually, this strategy means that the split criterion in RSF is different from the evaluation criterion of interest. In view of this discrepancy, we analyze the theoretical relationship between the two criteria and investigate whether a unified strategy that uses Harrell's C for both Node Splitting and evaluation is able to improve the performance of RSF. Based on simulation studies and the analysis of real-world data, we show that substantial performance gains are possible if the log-rank statistic is replaced by Harrell's C for Node Splitting in RSF. Our results also show that C-based Splitting is not superior to log-rank Splitting if the percentage of noise variables is high, a result which can be attributed to the more unbalanced splits that are generated by the log-rank statistic.

  • On the use of Harrell's C for clinical risk prediction via random survival forests
    arXiv: Machine Learning, 2015
    Co-Authors: Matthias Schmid, Marvin N. Wright, Andreas Ziegler
    Abstract:

    Random survival forests (RSF) are a powerful method for risk prediction of right-censored outcomes in biomedical research. RSF use the log-rank split criterion to form an ensemble of survival trees. The most common approach to evaluate the prediction accuracy of a RSF model is Harrell's concordance index for survival data ('C index'). Conceptually, this strategy implies that the split criterion in RSF is different from the evaluation criterion of interest. This discrepancy can be overcome by using Harrell's C for both Node Splitting and evaluation. We compare the difference between the two split criteria analytically and in simulation studies with respect to the preference of more unbalanced splits, termed end-cut preference (ECP). Specifically, we show that the log-rank statistic has a stronger ECP compared to the C index. In simulation studies and with the help of two medical data sets we demonstrate that the accuracy of RSF predictions, as measured by Harrell's C, can be improved if the log-rank statistic is replaced by the C index for Node Splitting. This is especially true in situations where the censoring rate or the fraction of informative continuous predictor variables is high. Conversely, log-rank Splitting is preferable in noisy scenarios. Both C-based and log-rank Splitting are implemented in the R~package ranger. We recommend Harrell's C as split criterion for use in smaller scale clinical studies and the log-rank split criterion for use in large-scale 'omics' studies.

Nicky J Welton - One of the best experts on this subject based on the ideXlab platform.

  • Automated generation of Node-Splitting models for assessment of inconsistency in network meta-analysis.
    Research synthesis methods, 2015
    Co-Authors: Gert Van Valkenhoef, Sofia Dias, A E Ades, Nicky J Welton
    Abstract:

    Network meta-analysis enables the simultaneous synthesis of a network of clinical trials comparing any number of treatments. Potential inconsistencies between estimates of relative treatment effects are an important concern, and several methods to detect inconsistency have been proposed. This paper is concerned with the Node-Splitting approach, which is particularly attractive because of its straightforward interpretation, contrasting estimates from both direct and indirect evidence. However, Node-Splitting analyses are labour-intensive because each comparison of interest requires a separate model. It would be advantageous if Node-Splitting models could be estimated automatically for all comparisons of interest. We present an unambiguous decision rule to choose which comparisons to split, and prove that it selects only comparisons in potentially inconsistent loops in the network, and that all potentially inconsistent loops in the network are investigated. Moreover, the decision rule circumvents problems with the parameterisation of multi-arm trials, ensuring that model generation is trivial in all cases. Thus, our methods eliminate most of the manual work involved in using the Node-Splitting approach, enabling the analyst to focus on interpreting the results.

  • Automated generation of Node- Splitting models for assessment of inconsistency in network meta-analysis Gert van Valkenhoef, a * Sofia Dias, b A. E. Ades b
    2015
    Co-Authors: Nicky J Welton
    Abstract:

    Network meta-analysis enables the simultaneous synthesis of a network of clinical trials comparing any number of treatments. Potential inconsistencies between estimates of relative treatment effects are an important concern, and several methods to detect inconsistency have been proposed. This paper is concerned with the Node-Splitting approach, which is particularly attractive because of its straightforward interpretation, contrasting estimates from both direct and indirect evidence. However, Node-Splitting analyses are labour-intensive because each comparison of interest requires a separate model. It would be advantageous if Node-Splitting models could be estimated automatically for all comparisons of interest. We present an unambiguous decision rule to choose which comparisons to split, and prove that it selects only comparisons in potentially inconsistent loops in the network, and that all potentially inconsistent loops in the network are investigated. Moreover, the decision rule circumvents problems with the parameterisation of multi-arm trials, ensuring that model generation is trivial in all cases. Thus, our methods eliminate most of the manual work involved in using the Node-Splitting approach, enabling the analyst to focus on interpreting the results. © 2015 The Authors Research Synthesis Methods Published by John Wiley & Sons Ltd.

  • automated generation of Node Splitting models for assessment of inconsistency in network meta analysis gert van valkenhoef a sofia dias b a e ades b
    2015
    Co-Authors: Nicky J Welton
    Abstract:

    Network meta-analysis enables the simultaneous synthesis of a network of clinical trials comparing any number of treatments. Potential inconsistencies between estimates of relative treatment effects are an important concern, and several methods to detect inconsistency have been proposed. This paper is concerned with the Node-Splitting approach, which is particularly attractive because of its straightforward interpretation, contrasting estimates from both direct and indirect evidence. However, Node-Splitting analyses are labour-intensive because each comparison of interest requires a separate model. It would be advantageous if Node-Splitting models could be estimated automatically for all comparisons of interest. We present an unambiguous decision rule to choose which comparisons to split, and prove that it selects only comparisons in potentially inconsistent loops in the network, and that all potentially inconsistent loops in the network are investigated. Moreover, the decision rule circumvents problems with the parameterisation of multi-arm trials, ensuring that model generation is trivial in all cases. Thus, our methods eliminate most of the manual work involved in using the Node-Splitting approach, enabling the analyst to focus on interpreting the results. © 2015 The Authors Research Synthesis Methods Published by John Wiley & Sons Ltd.

D Shin - One of the best experts on this subject based on the ideXlab platform.

  • Author name disambiguation using a graph model with Node Splitting and merging based on bibliographic information
    Scientometrics, 2014
    Co-Authors: D Shin, T. Kim, J Choi, Jungsun Kim
    Abstract:

    Author ambiguity mainly arises when several different authors express\ntheir names in the same way, generally known as the namesake problem,\nand also when the name of an author is expressed in many different ways,\nreferred to as the heteronymous name problem. These author ambiguity\nproblems have long been an obstacle to efficient information retrieval\nin digital libraries, causing incorrect identification of authors and\nimpeding correct classification of their publications. It is a\nnontrivial task to distinguish those authors, especially when there is\nvery limited information about them. In this paper, we propose a graph\nbased approach to author name disambiguation, where a graph model is\nconstructed using the co-author relations, and author ambiguity is\nresolved by graph operations such as vertex (or Node) Splitting and\nmerging based on the co-authorship. In our framework, called a Graph\nFramework for Author Disambiguation (GFAD), the namesake problem is\nsolved by Splitting an author vertex involved in multiple cycles of\ncoauthorship, and the heteronymous name problem is handled by merging\nmultiple author vertices having similar names if those vertices are\nconnected to a common vertex. Experiments were carried out with the real\nDBLP and Arnetminer collections and the performance of GFAD is compared\nwith three representative unsupervised author name disambiguation\nsystems. We confirm that GFAD shows better overall performance from the\nperspective of representative evaluation metrics. An additional\ncontribution is that we released the refined DBLP collection to the\npublic to facilitate organizing a performance benchmark for future\nsystems on author disambiguation.

R Anitha - One of the best experts on this subject based on the ideXlab platform.

  • a novel Node Splitting criteria for decision trees based on theil index
    International Conference on Neural Information Processing, 2012
    Co-Authors: Shina Sheen, R Anitha
    Abstract:

    The performance of detectors using decision trees can be improved by reducing the average height of the tree for faster detection. We propose a new attribute Splitting criteria for decision tree construction using the concept of Theil index. The Theil index is a statistic used to measure economic inequality. Results show a decrease in average height compared to the frequently used trees like ID3 and C4.5 using impurity measure as the Splitting criterion. Detection of malware using data mining techniques has been explored extensively. Techniques used for detecting malware based on structural features rely on being able to identify anomalies in the structure of executable files. These features might indicate that the file was created or infected to perform malicious activity. They are applied to a decision tree using Theil index as Splitting criterion for classification as malware or benign files.

  • ICONIP (2) - A novel Node Splitting criteria for decision trees based on theil index
    Neural Information Processing, 2012
    Co-Authors: Shina Sheen, R Anitha
    Abstract:

    The performance of detectors using decision trees can be improved by reducing the average height of the tree for faster detection. We propose a new attribute Splitting criteria for decision tree construction using the concept of Theil index. The Theil index is a statistic used to measure economic inequality. Results show a decrease in average height compared to the frequently used trees like ID3 and C4.5 using impurity measure as the Splitting criterion. Detection of malware using data mining techniques has been explored extensively. Techniques used for detecting malware based on structural features rely on being able to identify anomalies in the structure of executable files. These features might indicate that the file was created or infected to perform malicious activity. They are applied to a decision tree using Theil index as Splitting criterion for classification as malware or benign files.