The Experts below are selected from a list of 318 Experts worldwide ranked by ideXlab platform
Jiming Jiang - One of the best experts on this subject based on the ideXlab platform.
-
Sumca: simple, unified, Monte-Carlo-assisted approach to second-order Unbiased mean-squared prediction error estimation
Journal of the Royal Statistical Society: Series B (Statistical Methodology), 2020Co-Authors: Jiming Jiang, Mahmoud TorabiAbstract:We propose a simple, unified, Monte‐Carlo‐assisted approach (called ‘Sumca’) to second‐order Unbiased estimation of the mean‐squared prediction error (MSPE) of a small area Predictor. The MSPE estimator proposed is easy to derive, has a simple expression and applies to a broad range of Predictors that include the traditional empirical best Linear Unbiased Predictor, empirical best Predictor and post‐model‐selection empirical best Linear Unbiased Predictor and empirical best Predictor as special cases. Furthermore, the leading term of the MSPE estimator proposed is guaranteed positive; the lower order term corresponds to a bias correction, which can be evaluated via a Monte Carlo method. The computational burden for the Monte Carlo evaluation is much less, compared with other Monte‐Carlo‐based methods that have been used for producing second‐order Unbiased MSPE estimators, such as the double bootstrap and Monte Carlo jackknife. The Sumca estimator also has a nice stability feature. Theoretical and empirical results demonstrate properties and advantages of the Sumca estimator.
-
On Unbiasedness of the empirical BLUE and BLUP
Statistics & Probability Letters, 1999Co-Authors: Jiming JiangAbstract:Abstract Let y = Xβ + Zα + ϵ be a mixed Linear model, where β is a vector of fixed effects, α is a vector of random effects, and ϵ is a vector of errors. Kackar and Harville (1984) showed that the best Linear Unbiased estimator (BLUE) of β and the best Linear Unbiased Predictor (BLUP) of α remain Unbiased if the true variance components at which the BLUE and BLUP are computed are replaced by nonnegative, even and translation-invariant estimators, provided the expectations of the resulting empirical BLUE and BLUP exist. In this short note, we show that when there is a single random effect factor in the model, those expectations do exist.
-
a derivation of blup best Linear Unbiased Predictor
Statistics & Probability Letters, 1997Co-Authors: Jiming JiangAbstract:Abstract We show the best Linear Unbiased Predictor (BLUP) can be derived as the best Predictor (under normality) based on all error contrasts (i.e., transformation of data with mean 0). The result reveals an interesting connection between BLUP and REML—restricted or residual maximum likelihood—estimates.
-
A derivation of BLUP—Best Linear Unbiased Predictor
Statistics & Probability Letters, 1997Co-Authors: Jiming JiangAbstract:Abstract We show the best Linear Unbiased Predictor (BLUP) can be derived as the best Predictor (under normality) based on all error contrasts (i.e., transformation of data with mean 0). The result reveals an interesting connection between BLUP and REML—restricted or residual maximum likelihood—estimates.
I Misztal - One of the best experts on this subject based on the ideXlab platform.
-
using single step genomic best Linear Unbiased Predictor to enhance the mitigation of seasonal losses due to heat stress in pigs
Journal of Animal Science, 2016Co-Authors: B O Fragomeni, Daniela Lourenco, S Tsuruta, H L Bradford, Kent Gray, Yijian Huang, I MisztalAbstract:The purposes of this study were to analyze the impact of seasonal losses due to heat stress in pigs from different breeds raised in different environments and to evaluate the accuracy improvement from adding genomic information to genetic evaluations. Data were available for 2 different swine populations: purebred Duroc animals raised in Texas and North Carolina and commercial crosses of Duroc and F females (Landrace × Large White) raised in Missouri and North Carolina; pedigrees provided links for animals from different states. Pedigree information was available for 553,442 animals, of which 8,232 pure breeds were genotyped. Traits were BW at 170 d for purebred animals and HCW for crossbred animals. Analyses were done with an animal model as either single- or 2-trait models using phenotypes measured in different states as separate traits. Additionally, reaction norm models were fitted for 1 or 2 traits using heat load index as a covariable. Heat load was calculated as temperature-humidity index greater than 70 and was averaged over 30 d prior to data collection. Variance components were estimated with average information REML, and EBV and genomic EBV (GEBV) with BLUP or single-step genomic BLUP (ssGBLUP). Validation was assessed for 146 genotyped sires with progeny in the last generation. Accuracy was calculated as a correlation between EBV and GEBV using reduced data (all animals, except the last generation) and using complete data. Heritability estimates for purebred animals were similar across states (varying from 0.23 to 0.26), and reaction norm models did not show evidence of a heat stress effect. Genetic correlations between states for heat loads were always strong (>0.91). For crossbred animals, no differences in heritability were found in single- or 2-trait analysis (from 0.17 to 0.18), and genetic correlations between states were moderate (0.43). In the reaction norm for crossbreeds, heritabilities ranged from 0.15 to 0.30 and genetic correlations between heat loads were as weak as 0.36, with heat load ranging from 0 to 12. Accuracies with ssGBLUP were, on average, 25% greater than with BLUP. Accuracies were greater in 2-trait reaction norm models and at extreme heat load values. Impacts of seasonality are evident only for crossbred animals. Genomic information can help producers mitigate heat stress in swine by identifying superior sires that are more resistant to heat stress.
-
crossbreed evaluations in single step genomic best Linear Unbiased Predictor using adjusted realized relationship matrices
Journal of Animal Science, 2016Co-Authors: Daniela Lourenco, S Tsuruta, B O Fragomeni, Chingyi Chen, W O Herring, I MisztalAbstract:Combining purebreed and crossbreed information is beneficial for genetic evaluation of some livestock species. Genetic evaluations can use relationships based on genomic information, relying on allele frequencies that are breed specific. Single-step genomic BLUP (ssGBLUP) does not account for different allele frequencies, which could limit the genetic gain in crossbreed evaluations. In this study, we tested the performance of different breed-specific genomic relationship matrices () in ssGBLUP for crossbreed evaluations; we also tested the importance of genotyping crossbred animals. Genotypes were available for purebreeds (AA and BB) and crossbreeds (F) in simulated and real pig populations. The number of genotyped animals was, on average, 4,315 for the simulated population and 15,798 for the real population. Cross-validation was performed on 1,200 and 3,117 F animals in the simulated and real populations, respectively. Simulated scenarios were under no artificial selection, mass selection, or BLUP selection. Two genomic relationship matrices were constructed based on breed-specific allele frequencies: 1) , a genomic relationship matrix centered by breed-specific allele frequencies, and 2) , a genomic relationship matrix centered and scaled by breed-specific allele frequencies. All (the across-breed genomic relationship matrix), , and were also tuned to account for selective genotyping. Using breed-specific allele frequencies reduced the number of negative relationships between 2 purebreeds, pulling the average closer to 0, as in the pedigree-based relationship matrix. For simulated populations that included mass selection, genomic EBV (GEBV) in F, when using and , were, on average, 10% more accurate than ; however, after tuning to account for selective genotyping, provided the same accuracy as for breed-specific genomic relationship matrices. For the real population, accuracies for litter size in F were 0.62 for , , and , and tuning had no impact on accuracy, except for , which was 1 percentage point less accurate. Accuracy of GEBV for number of stillborns in F1 was 0.5 for all tested genomic relationship matrices with no changes after tuning. We observed that genotyping F increased accuracies of GEBV for the same animals by up to 39% compared with having genotypes for only AA and BB. In crossbreed evaluations, accounting for breed-specific allele frequencies promoted changes in G that were not influential enough to improve accuracy of GEBV. Therefore, the best performance of ssGBLUP for crossbreed evaluations requires genotypes for pure- and crossbreeds and no breed-specific adjustments in the realized relationship matrix.
-
hot topic use of genomic recursions in single step genomic best Linear Unbiased Predictor blup with a large number of genotypes
Journal of Dairy Science, 2015Co-Authors: B O Fragomeni, Daniela Lourenco, S Tsuruta, Yutaka Masuda, Andres Legarra, I Aguilar, T J Lawlor, I MisztalAbstract:Abstract The purpose of this study was to evaluate the accuracy of genomic selection in single-step genomic BLUP (ssGBLUP) when the inverse of the genomic relationship matrix ( G ) is derived by the "algorithm for proven and young animals" (APY). This algorithm implements genomic recursions on a subset of "proven" animals. Only a relationship matrix for animals treated as "proven" needs to be inverted, and the extra costs of adding animals treated as "young" are Linear. Analyses involved 10,102,702 final scores on 6,930,618 Holstein cows. Final score, which is a composite of type traits, is popular trait in the United States and was easily available for this study. A total of 100,000 animals with genotypes were used in the analyses and included 23,000 sires (16,000 with >5 progeny), 27,000 cows, and 50,000 young animals. Genomic EBV (GEBV) were calculated with a regular inverse of G , and with the G inverse approximated by APY. Animals in the proven subset included only sires (23,000), sires + cows (50,000), only cows (27,000), or sires with >5 progeny (16,000). The correlations of GEBV with APY and regular GEBV for young genotyped animals were 0.994, 0.995, 0.992, and 0.992, respectively Later, animals in the proven subset were randomly sampled from all genotyped animals in sets of 2,000, 5,000, 10,000, 15,000, and 20,000; each sample was replicated 4 times. Respective correlations were 0.97 (5,000 sample), 0.98 (10,000 sample), and 0.99 (20,000 sample), with minimal difference between samples of the same size. Genomic EBV with APY were accurate when the number of animals used in the subset is between 10,000 and 20,000, with little difference between the ways of creating the subset. Due to the approximately Linear cost of APY, ssGBLUP with APY could support any number of genotyped animals without affecting accuracy.
B O Fragomeni - One of the best experts on this subject based on the ideXlab platform.
-
using single step genomic best Linear Unbiased Predictor to enhance the mitigation of seasonal losses due to heat stress in pigs
Journal of Animal Science, 2016Co-Authors: B O Fragomeni, Daniela Lourenco, S Tsuruta, H L Bradford, Kent Gray, Yijian Huang, I MisztalAbstract:The purposes of this study were to analyze the impact of seasonal losses due to heat stress in pigs from different breeds raised in different environments and to evaluate the accuracy improvement from adding genomic information to genetic evaluations. Data were available for 2 different swine populations: purebred Duroc animals raised in Texas and North Carolina and commercial crosses of Duroc and F females (Landrace × Large White) raised in Missouri and North Carolina; pedigrees provided links for animals from different states. Pedigree information was available for 553,442 animals, of which 8,232 pure breeds were genotyped. Traits were BW at 170 d for purebred animals and HCW for crossbred animals. Analyses were done with an animal model as either single- or 2-trait models using phenotypes measured in different states as separate traits. Additionally, reaction norm models were fitted for 1 or 2 traits using heat load index as a covariable. Heat load was calculated as temperature-humidity index greater than 70 and was averaged over 30 d prior to data collection. Variance components were estimated with average information REML, and EBV and genomic EBV (GEBV) with BLUP or single-step genomic BLUP (ssGBLUP). Validation was assessed for 146 genotyped sires with progeny in the last generation. Accuracy was calculated as a correlation between EBV and GEBV using reduced data (all animals, except the last generation) and using complete data. Heritability estimates for purebred animals were similar across states (varying from 0.23 to 0.26), and reaction norm models did not show evidence of a heat stress effect. Genetic correlations between states for heat loads were always strong (>0.91). For crossbred animals, no differences in heritability were found in single- or 2-trait analysis (from 0.17 to 0.18), and genetic correlations between states were moderate (0.43). In the reaction norm for crossbreeds, heritabilities ranged from 0.15 to 0.30 and genetic correlations between heat loads were as weak as 0.36, with heat load ranging from 0 to 12. Accuracies with ssGBLUP were, on average, 25% greater than with BLUP. Accuracies were greater in 2-trait reaction norm models and at extreme heat load values. Impacts of seasonality are evident only for crossbred animals. Genomic information can help producers mitigate heat stress in swine by identifying superior sires that are more resistant to heat stress.
-
crossbreed evaluations in single step genomic best Linear Unbiased Predictor using adjusted realized relationship matrices
Journal of Animal Science, 2016Co-Authors: Daniela Lourenco, S Tsuruta, B O Fragomeni, Chingyi Chen, W O Herring, I MisztalAbstract:Combining purebreed and crossbreed information is beneficial for genetic evaluation of some livestock species. Genetic evaluations can use relationships based on genomic information, relying on allele frequencies that are breed specific. Single-step genomic BLUP (ssGBLUP) does not account for different allele frequencies, which could limit the genetic gain in crossbreed evaluations. In this study, we tested the performance of different breed-specific genomic relationship matrices () in ssGBLUP for crossbreed evaluations; we also tested the importance of genotyping crossbred animals. Genotypes were available for purebreeds (AA and BB) and crossbreeds (F) in simulated and real pig populations. The number of genotyped animals was, on average, 4,315 for the simulated population and 15,798 for the real population. Cross-validation was performed on 1,200 and 3,117 F animals in the simulated and real populations, respectively. Simulated scenarios were under no artificial selection, mass selection, or BLUP selection. Two genomic relationship matrices were constructed based on breed-specific allele frequencies: 1) , a genomic relationship matrix centered by breed-specific allele frequencies, and 2) , a genomic relationship matrix centered and scaled by breed-specific allele frequencies. All (the across-breed genomic relationship matrix), , and were also tuned to account for selective genotyping. Using breed-specific allele frequencies reduced the number of negative relationships between 2 purebreeds, pulling the average closer to 0, as in the pedigree-based relationship matrix. For simulated populations that included mass selection, genomic EBV (GEBV) in F, when using and , were, on average, 10% more accurate than ; however, after tuning to account for selective genotyping, provided the same accuracy as for breed-specific genomic relationship matrices. For the real population, accuracies for litter size in F were 0.62 for , , and , and tuning had no impact on accuracy, except for , which was 1 percentage point less accurate. Accuracy of GEBV for number of stillborns in F1 was 0.5 for all tested genomic relationship matrices with no changes after tuning. We observed that genotyping F increased accuracies of GEBV for the same animals by up to 39% compared with having genotypes for only AA and BB. In crossbreed evaluations, accounting for breed-specific allele frequencies promoted changes in G that were not influential enough to improve accuracy of GEBV. Therefore, the best performance of ssGBLUP for crossbreed evaluations requires genotypes for pure- and crossbreeds and no breed-specific adjustments in the realized relationship matrix.
-
hot topic use of genomic recursions in single step genomic best Linear Unbiased Predictor blup with a large number of genotypes
Journal of Dairy Science, 2015Co-Authors: B O Fragomeni, Daniela Lourenco, S Tsuruta, Yutaka Masuda, Andres Legarra, I Aguilar, T J Lawlor, I MisztalAbstract:Abstract The purpose of this study was to evaluate the accuracy of genomic selection in single-step genomic BLUP (ssGBLUP) when the inverse of the genomic relationship matrix ( G ) is derived by the "algorithm for proven and young animals" (APY). This algorithm implements genomic recursions on a subset of "proven" animals. Only a relationship matrix for animals treated as "proven" needs to be inverted, and the extra costs of adding animals treated as "young" are Linear. Analyses involved 10,102,702 final scores on 6,930,618 Holstein cows. Final score, which is a composite of type traits, is popular trait in the United States and was easily available for this study. A total of 100,000 animals with genotypes were used in the analyses and included 23,000 sires (16,000 with >5 progeny), 27,000 cows, and 50,000 young animals. Genomic EBV (GEBV) were calculated with a regular inverse of G , and with the G inverse approximated by APY. Animals in the proven subset included only sires (23,000), sires + cows (50,000), only cows (27,000), or sires with >5 progeny (16,000). The correlations of GEBV with APY and regular GEBV for young genotyped animals were 0.994, 0.995, 0.992, and 0.992, respectively Later, animals in the proven subset were randomly sampled from all genotyped animals in sets of 2,000, 5,000, 10,000, 15,000, and 20,000; each sample was replicated 4 times. Respective correlations were 0.97 (5,000 sample), 0.98 (10,000 sample), and 0.99 (20,000 sample), with minimal difference between samples of the same size. Genomic EBV with APY were accurate when the number of animals used in the subset is between 10,000 and 20,000, with little difference between the ways of creating the subset. Due to the approximately Linear cost of APY, ssGBLUP with APY could support any number of genotyped animals without affecting accuracy.
-
Genetic evaluation using single-step genomic best Linear Unbiased Predictor in American Angus
Journal of Animal Science, 2015Co-Authors: D.a.l Lourenco, B O Fragomeni, Shogo Tsuruta, Yutaka Masuda, Ignacio Aguilar, Andres Legarra, J.k. Bertrand, T.s. Amen, L. Wang, D.w. MoserAbstract:Predictive ability of genomic EBV when using single-step genomic BLUP (ssGBLUP) in Angus cattle was investigated. Over 6 million records were available on birth weight (BiW) and weaning weight (WW), almost 3.4 million on postweaning gain (PWG), and over 1.3 million on calving ease (CE). Genomic information was available on, at most, 51,883 animals, which included high and low EBV accuracy animals. Traditional EBV was computed by BLUP and genomic EBV by ssGBLUP and indirect prediction based on SNP effects was derived from ssGBLUP; SNP effects were calculated based on the following reference populations: ref_2k (contains top bulls and top cows that had an EBV accuracy for BiW ≥0.85), ref_8k (contains all parents that were genotyped), and ref_33k (contains all genotyped animals born up to 2012). Indirect prediction was obtained as direct genomic value (DGV) or as an index of DGV and parent average (PA). Additionally, runs with ssGBLUP used the inverse of the genomic relationship matrix calculated by an algorithm for proven and young animals (APY) that uses recursions on a small subset of reference animals. An extra reference subset included 3,872 genotyped parents of genotyped animals (ref_4k). Cross-validation was used to assess predictive ability on a validation population of 18,721 animals born in 2013. Computations for growth traits used multiple-trait Linear model and, for CE, a bivariate CE-BiW threshold-Linear model. With BLUP, predictivities were 0.29, 0.34, 0.23, and 0.12 for BiW, WW, PWG, and CE, respectively. With ssGBLUP and ref_2k, predictivities were 0.34, 0.35, 0.27, and 0.13 for BiW, WW, PWG, and CE, respectively, and with ssGBLUP and ref_33k, predictivities were 0.39, 0.38, 0.29, and 0.13 for BiW, WW, PWG, and CE, respectively. Low predictivity for CE was due to low incidence rate of difficult calving. Indirect predictions with ref_33k were as accurate as with full ssGBLUP. Using the APY and recursions on ref_4k gave 88% gains of full ssGBLUP and using the APY and recursions on ref_8k gave 97% gains of full ssGBLUP. Genomic evaluation in beef cattle with ssGBLUP is feasible while keeping the models (maternal, multiple trait, and threshold) already used in regular BLUP. Gains in predictivity are dependent on the composition of the reference population. Indirect predictions via SNP effects derived from ssGBLUP allow for accurate genomic predictions on young animals, with no advantage of including PA in the index if the reference population is large. With the APY conditioning on about 10,000 reference animals, ssGBLUP is potentially applicable to a large number of genotyped animals without compromising predictive ability.
Nicola Salvati - One of the best experts on this subject based on the ideXlab platform.
-
Small Area Estimation in Practice: An Application to Agricultural Business Survey Data
2012Co-Authors: Nikos Tzavidis, Ray Chambers, Nicola Salvati, Hukum ChandraAbstract:This paper describes an application of small area estimation (SAbl to agricultural business survey data. Both well known small area estimators, such as the empirical best Linear Unbiased Predictor (EBLUP), and more recently proposed small area estimators, for example, tile M-quanlile, the robust EBLUP aml!he Model Ansed Direct estimators arc considered. Mean squared error estimation is discussed. Using a real agricultural business survey dataset, we place emphasis on model diagnostics for specifying the small area working model, on diagnostic measures for validating the reliability of direct and indirect (modelbased) small area e.~timators i1nd on providing practical guidelines to the prospective user of small area estimation techniques.
-
Small area estimation under spatial nonstationarity
Computational Statistics & Data Analysis, 2012Co-Authors: Hukum Chandra, Nicola Salvati, Ray Chambers, Nikos TzavidisAbstract:A geographical weighted empirical best Linear Unbiased Predictor (GWEBLUP) for a small area average is proposed, and an estimator of its conditional mean squared error is developed. The popular empirical best Linear Unbiased Predictor under the Linear mixed model is obtained as a special case of the GWEBLUP. Empirical results using both model-based and design-based simulations, with the latter based on two real data sets, show that the GWEBLUP Predictor can lead to efficiency gains when spatial nonstationarity is present in the data. A practical gain from using the GWEBLUP is in small area estimation for out of sample areas. In this case the efficient use of geographical information can potentially improve upon conventional synthetic estimation.
-
Small area estimation: the EBLUP estimator based on spatially correlated random area effects
Statistical Methods and Applications, 2008Co-Authors: Monica Pratesi, Nicola SalvatiAbstract:This paper deals with small area indirect estimators under area level random effect models when only area level data are available and the random effects are correlated. The performance of the Spatial Empirical Best Linear Unbiased Predictor (SEBLUP) is explored with a Monte Carlo simulation study on lattice data and it is applied to the results of the sample survey on Life Conditions in Tuscany (Italy). The mean squared error (MSE) problem is discussed illustrating the MSE estimator in comparison with the MSE of the empirical sampling distribution of SEBLUP estimator. A clear tendency in our empirical findings is that the introduction of spatially correlated random area effects reduce both the variance and the bias of the EBLUP estimator. Despite some residual bias, the coverage rate of our confidence intervals comes close to a nominal 95%.
-
Bootstrap for estimating the mean squared error of the spatial EBLUP
2007Co-Authors: Monica Pratesi, Nicola Salvati, Isabel MolinaAbstract:This work assumes that the small area quantities of interest follow a Fay-Herriot model with spatially correlated random area effects. Under this model, parametric and nonparametric bootstrap procedures are proposed for estimating the mean squared error of the EBLUP (Empirical Best Linear Unbiased Predictor). A simulation study compares the bootstrap estimates with an asymptotic analytical approximation and studies the robustness to non-normality. Finally, two applications with real data are described.
-
Small Area Estimation Using Spatial Information. The Rathbun Lake Watershed Case Study
2003Co-Authors: Alessandra Petrucci, Nicola Salvati, G. ParentiAbstract:The paper describes an application of a modified small area estimator to the data collected in the Rathbun Lake Watershed in Iowa (USA). Opsomer et al. (2003) estimated the average erosion per acre for 61 sub-watersheds within the study region using an empirical best Linear Unbiased Predictor (EBLUP) and a composite estimator. The proposed methodology considers an EBLUP estimator with spatially correlated error taking into account the information provided by neighboring areas.
Tapabrata Maiti - One of the best experts on this subject based on the ideXlab platform.
-
Mean-squared error estimation in transformed Fay-Herriot models
Journal of the Royal Statistical Society: Series B (Statistical Methodology), 2006Co-Authors: Eric V. Slud, Tapabrata MaitiAbstract:Summary. The problem of accurately estimating the mean-squared error of small area estimators within a Fay–Herriot normal error model is studied theoretically in the common setting where the model is fitted to a logarithmically transformed response variable. For bias-corrected empirical best Linear Unbiased Predictor small area point estimators, mean-squared error formulae and estimators are provided, with biases of smaller order than the reciprocal of the number of small areas. The performance of these mean-squared error estimators is illustrated by a simulation study and a real data example relating to the county level estimation of child poverty rates in the US Census Bureau’s on-going ‘Small area income and poverty estimation’ project.