Skip to main content
Software

SNP-GWAS: a flexible genome-wide association study software adapted for very large designs

Authors
  • Thierry Tribout (Université Paris-Saclay)
  • Didier Boichard (Université Paris-Saclay)

Abstract

Genome-wide association studies (GWAS) at the whole genome sequence level are the method of choice for mapping candidate causal variants underlying QTLs. Power and resolution are highly dependent on the size of the population analysed and it is therefore strongly advised to analyse very large populations. To take into account the population structure and avoid potential confounding and false positives, most approaches include a polygenic effect (g). The genomic relationship matrix G built from SNP genotypes is dense and makes the computations very long or even impossible when the number of individuals is large (e.g. more than 40,000). In animal breeding, the fastGWAS approach (which sets to zero the small terms of G) is not appropriate due to the high average relationship between individuals. Here we propose to replace g by Ms, i.e. by the sum of the SNP effects s multiplied by the centred genotypes M, for a set of a few tens of thousands of markers distributed across the genome and chosen for their informativeness. The advantage of this option is that the size of the M'M matrix is equal to the number of SNPs, making the size of the equation system constant regardless of the number of individuals analysed. In addition, our software includes four desirable features which are not always proposed simultaneously. Records can be weighed by the inverse of their residual variance. If phenotypes have different accuracies, this method increases power and reduces false positive rates. When sequence variants are imputed from chip genotypes, their imputation accuracy is less than 1, and the software can use dosages, i.e. expected genotypes, instead of genotypes assumed to be known, again to maximise power and reduce false positive rates. We also propose to optionally estimate dominance in addition to additive effect. Finally, to maximise power, it is generally recommended to remove the markers on the same chromosome as the tested SNP, using the so-called LOCO procedure. Our software allows to exclude the whole chromosome or only variants in a user-defined segment surrounding the tested SNP ("LOSO"). All these features are included in the proposed software. The software has been conceived to perform single-trait GWAS at sequence level, i.e. for a very large number of biallelic variants. It analyses the chromosomes separately, by parallelizing analyses of groups of variants of a same chromosome. It is developed in Fortran 90, and makes extensive use of optimized matrix calculation functions from the MKL library. In addition, it includes a number of features to minimize computations. A large-scale example is showed with the analysis of age at first calving of 145,000 Montbéliarde cows and 900,000 variants of chromosome 21. The software will be made publicly available before July 2026.

Keywords: 2026

How to Cite:

Tribout, T. & Boichard, D., (2026) “SNP-GWAS: a flexible genome-wide association study software adapted for very large designs”, World Congress on Genetics Applied to Livestock Production Digital Archive 2026(1): 2285960. doi: https://doi.org/10.31274/wcgalp.23824

Rights: 1

Downloads:
Download PDF
View PDF

75 Views

19 Downloads

Published on
2026-02-25

Peer Reviewed