Variance components estimation with Monte-Carlo single-step genomic REML
Abstract
As phenotypic and genomic information accumulates, estimating variance components (VC) becomes computationally infeasible because restricted maximum likelihood (REML) methods require the inverse of the mixed model equations (MME) to compute the trace term of the first derivatives of the restricted log-likelihood. Monte Carlo REML (MC-REML) employs a Monte Carlo estimator of the trace to avoid inverting the MME by simulating the data and solving the MME for each sample. Previously, MC-REML did not consider genomic information for the simulation process, making it unsuitable for single-step genomic BLUP (ssGBLUP) models. We developed Monte Carlo REML for ssGBLUP models, which we called it as Monte Carlo single-step genomic REML (MC-ssGREML). This study aims to present the theoretical formulation, an application example of MC-ssGREML, and early studies on improving the method's convergence robustness. We derived formulas to efficiently simulate the ssGBLUP data-generating process, including marker effects, residual polygenic effects, J factors or metafounders, and imputation errors. We simultaneously simulate correlated random effects and traits, enhancing computing efficiency. Convergence problems with MC-REML were experienced in the past due to Monte Carlo errors. We proposed a new convergence criterion consisting of the coefficient of variation of the five previous runs of the approximated restricted likelihood (CVlogL). This convergence criteria may be limited by the MC error near convergence when the sample size is insufficient. We implemented an adaptive sampling scheme that increased the sample count by five whenever the norm of the scaled first derivatives of the restricted log-likelihood Δc fluctuates near convergence. We tested MC-ssGREML for a beef cattle production model with maternal effects. Three traits were analyzed: birth weight, weaning weight, and post-weaning gain. We compared the exact AI-REML with AI-MC-ssGREML on a small dataset of 100,000 animals, of which 10,000 were genotyped. The model had 14 VCs to estimate. We used five samples to estimate the trace terms, and only one sample for the likelihood approximation. A large dataset of 7.4 million animals in the pedigree, from which 330,000 were genotyped, was used to assess the computing efficiency of MC-ssGREML in large-scale applications. For the small dataset, we did not observe significant differences between the exact AI-REML and MC-ssGREML approaches, with running times of 154 and 22 hours, respectively. The computing time for the large dataset was five days, with a memory usage of only 53 GB. Adaptive sampling based solely on Δc did not reduce MC error or increase robustness of MC-ssGREML convergence. Nevertheless, the current convergence criteria, CVlogL, complete the estimation without significant differences between AI-REML and MC-ssGREML estimates. Overall, MC-ssGREML is a useful framework for an efficient VC estimation in large, complex genomic datasets.
Keywords: 2026
How to Cite:
Bernardes, M., Bermann, M., Álvarez-Munera, A., Legarra, A., Misztal, I. & Lourenco, D., (2026) “Variance components estimation with Monte-Carlo single-step genomic REML”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2284216. doi: https://doi.org/10.31274/wcgalp.23558
Rights: 1
Downloads:
Download PDF
View PDF
151 Views
29 Downloads