Approximating GWAS with single-step SNPBLUP
Abstract
Genome-wide association studies (GWAS) are an indispensable tool for identifying loci associated with complex traits. GWAS software relying on mixed linear model association (mlma) analysis (e.g. GCTA) are considered the gold standard, yet more efficient software packages (e.g. BOLT-LMM, LDAK-KVIK) relying on approximations have been developed for balanced human populations, but lead to high genomic inflation ( >2) in related livestock populations. Moreover, in livestock, the increasing availability of large datasets with many unbalanced effects, complex mixed linear models, and the presence of ungenotyped and genotyped animals presents significant (computational) challenges. This study aims to address these challenges by approximating GWAS using single-step SNP best linear unbiased prediction (ssSNPBLUP). The proposed approximate GWAS approach includes two steps. First, SNP effects are estimated by solving ssSNPBLUP with a preconditioned conjugate gradient method. Second, prediction error variances for all SNP effects, needed for computing p-values, are approximated by the diagonal elements of the inverse of the coefficient matrix of a SNPBLUP model using information of both genotyped and ungenotyped animals. For efficiency, a sliding-window approach was used assuming linkage disequilibrium between SNPs more than 50,000 SNPs apart is zero, resulting in multiple inversions of coefficient matrices of size equal to 50,000, instead of one single inversion of a matrix encompassing all SNPs. To assess the correctness and performance of our approach, we analysed an imputed whole-genome sequence (WGS) dataset comprising 100,968 animals and 19,314,887 SNPs with minor allele frequency >0.01. Of these, 81,484 animals were phenotyped, and pedigree information was included. We validated our approach implemented in MiXBLUP against the mlma approach implemented in GCTA. Using a subset of 44K SNPs, the Pearson correlation between log-transformed p-values of both approaches was 0.999. Using a LD-pruned WGS data (4.8M SNPs), MiXBLUP closely approximated the GCTA results: 7 out of 10 lead variants identified by MiXBLUP were within 500Kb of a GCTA lead variant, 2 were within 3 Mbp, and 1 was uniquely detected by MiXBLUP. Within-chromosome Pearson correlations between log-transformed p-values ranged from 0.88 to 0.93. Genomic inflation factors were below 1. Regarding performance, MiXBLUP required 34 and 197 hours of wall clock time for the 4.8M and 19M SNPs datasets, respectively, with maximum RAM usage of 407 and 902 GB. The approximate GWAS method was further evaluated using a routine Topigs Norsvin dataset comprising 1.8M phenotype records for 10 production traits and 1.2M genotypes of approximately 20K SNPs, requiring 3 hours and 174 GB of RAM. Overall, these results demonstrate that the proposed method is practical and efficient for GWAS of large, complex datasets with unbalanced designs and mixtures of genotyped and ungenotyped animals, as commonly encountered in animal breeding.
Keywords: 2026
How to Cite:
Bouwman, A., Huisman, J., Stevens, T. & Vandenplas, J., (2026) “Approximating GWAS with single-step SNPBLUP”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285042. doi: https://doi.org/10.31274/wcgalp.23629
Rights: 1
Downloads:
Download PDF
View PDF
53 Views
14 Downloads