Skip to main content
Estimation & Prediction

Bayesian estimation of the relationships across metafounders or base populations

Authors
  • Cristina Álvarez-Múnera (University of Georgia)
  • Matias Bermann (University of Georgia)
  • Andres Legarra (University of Georgia)
  • Ignacy Misztal (University of Georgia)
  • Daniela Lourenco (University of Georgia)

Abstract

Accurate estimation of the relationship matrix among metafounders or base populations (Γ) is essential to maintain consistency between pedigree and genomic information and to avoid bias in genetic evaluations. The Γ can be derived from base population allele frequencies, which are not directly observable. In recent years, several methods such as generalized least squares (GLS), maximum likelihood (ML), and pseudo-expectation-maximization (pseudo-EM) have been proposed. While GLS and pseudo-EM are biased by construction and do not quantify the uncertainty associated with Γ estimates, ML can only be applied to specific population structures.To address these limitations, we developed a Bayesian approach based on the Metropolis-Hastings (MH) algorithm to estimate Γ through sampling of the base population allele frequencies. Each frequency was modeled as an independent Beta distribution, with shape parameters obtained from prior estimates (GLS or current population allele frequencies). Two sampling strategies were implemented: an independent sampler, which draws proposals directly from the initial Beta distributions, and an adaptive sampler, which updates the proposal parameters based on previously accepted samples to improve mixing and convergence.Analyses were performed using a simulated dataset including a pedigree of 84,202 animals, genotypes for 2,117 founders, and 40,000 SNPs, with 2 metafounders defined through prior allele frequencies. Both the adaptive and the independent samplers were run for 10,000 iterations. The independent sampler had a very low acceptance rate (∼0.18%) and a small effective sample size, although posterior means were close to the simulated values. However, posterior standard deviations were one order of magnitude smaller than the true parameters. The adaptive sampler achieved a higher acceptance rate (∼11%) and larger ESS, with similar posterior means and posterior standard deviations closer to the true parameters. Both samplers recovered the simulated values, but the adaptive sampler was more efficient.The proposed approaches do not have theoretical approximations like the likelihood truncation in pseudo-EM and provide posterior standard errors and uncertainty measures for each Γ element. Although computationally intensive for large-scale datasets, the methods provides a coherent statistical framework for research, method comparison, and formal assessment of variability in genomic relationships among base populations. Given its superior acceptance rate, mixing, and effective sample size, the adaptive sampler should be preferred for routine estimation of Γ, whereas the independent sampler may be useful when strong prior information is available or for exploratory analyses.

Keywords: 2026

How to Cite:

Álvarez-Múnera, C., Bermann, M., Legarra, A., Misztal, I. & Lourenco, D., (2026) “Bayesian estimation of the relationships across metafounders or base populations”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2286677. doi: https://doi.org/10.31274/wcgalp.24094

Rights: 1

Downloads:
Download PDF
View PDF

74 Views

18 Downloads

Published on
2026-02-26

Peer Reviewed