Skip to main content
Estimation & Prediction

Fitting linear mixed models to millions of whole genomes using the ancestral recombination graph

Authors
  • Gregor Gorjanc (The University of Edinburgh)
  • Brieuc Lehmann (University College London)
  • Hanbin Lee (University of Michigan)
  • Nathaniel S. Pope (University of Oregon)
  • Jerome Kelleher (University of Oxford)
  • Peter L. Ralph (University of Oregon)

Abstract

Ancestral recombination graphs (ARGs) are attractive for genomic analysis because they encode realised genetic relatedness between individuals as a function of past DNA branching, recombination, and mutation events. Recent ARG research has delivered tree sequence data structure for efficiently storing large genome datatsets, methods for inferring ARGs from observed genomes, fast algorithms to analyze ARGs, and applications of these advances with tens to hundreds of thousands of genomes. Because of their expresivity and scalability, ARGs provide a foundation for fitting linear mixed models (LMMs) to large genome datasets. We have developed a generative model of complex traits with additive effects on an ARG (ARG-LMM), associated ARG genetic relatedness matrix (ARG-GRM), scalable matrix-vector product algorithms with ARG-GRM without forming the matrix, and open-source software tskit and tslmm to fit the ARG-LMM to observed data. The ARG-LMM is defined on ARG edges, representing genome regions (haplotypes) passed between ancestors and descendants. These edges can represent within-pedigree inheritance between parents and their children or more distant inheritance within or even between (sub-)populations (from the perspective of most recent common ancestors of descendants). The associated edge effects represent the cumulative effect of mutations occurring on the haplotypes. The summation of edge effects from most recent common ancestors to descendants produces haplotype values. The summation of haplotype values carried by each individual produces individual genetic values. These summations are performed efficiently using tree sequence traversal algorithms. The ARG-LMM is parameterized by a variance component proportional to mutational variance under the infinite-sites model of mutation from population genetics and additive trait model from quantitative genetics. To fit the ARG-LMM we developed efficient algorithms for the average information restricted maximum likelihood (AI-REML) to estimate variance components and the best linear unbiased prediction (BLUP) to estimate genetic values. The efficiency is obtained by leveraging the tree sequence structure of an ARG and by using randomised linear algebra and conjugate gradient method with randomised Nyström preconditioning. The same algorithms can be used to estimate haplotype and mutation effects, meaning that the same framework can be used for genomic prediction and genome-wide association studies, while conditioning on the observed genomic relatedness between individuals within and across different populations. These populations can represent different families, breeds, or even sub-species. The ARG-LMM extends to structural variants encoded in an ARG, provided that such ARG can be inferred from observed genomic data. Our results show that the algorithms scale nearly linearly in the number of individuals, that we recover true variance components and competitive accuracy of estimated genetic values in simulations, both with true and inferred ARGs. This work will enable quantitative genetic analysis of complex traits in datasets with millions of individuals with whole-genome sequence data.

Keywords: 2026

How to Cite:

Gorjanc, G., Lehmann, B., Lee, H., Pope, N., Kelleher, J. & Ralph, P. L., (2026) “Fitting linear mixed models to millions of whole genomes using the ancestral recombination graph”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2286304. doi: https://doi.org/10.31274/wcgalp.24582

Rights: 1

Downloads:
Download PDF
View PDF

83 Views

17 Downloads

Published on
2026-02-25

Peer Reviewed