Abstract
Genetic variation among human populations is relatively small, falling below the threshold commonly used to define biological subspecies, prompting a reconsideration of traditional racial classifications. The term Biogeographical Ancestry (BGA) was introduced as a more objective and biologically grounded alternative, describing an individual’s ancestral origin in relation to contemporary population structure. BGA has proven valuable for forensic genetic profiling, studies of disease susceptibility, and investigations into human evolutionary history and migration patterns.Current BGA inference methods rely on ancestry informative markers (AIMs). While using a small set of genetic markers offers practical advantages, this approach has several critical limitations, including limited geographic coverage due to insufficient representation of global genetic diversity, and reduced inference accuracy arising from oversimplified modelling assumptions, such as linkage equilibrium and absence of recent admixture, which do not fully reflect biological complexities, especially in admixed or geographically intermediate populations.
This study addresses the problem of BGA inference by directly utilising genome-wide single nucleotide polymorphism (SNP) variation. The aim of this research is to develop a non-AIM-based BGA inference framework capable of delivering stable and accurate performance for both non-admixed and admixed individuals across seven global intra-continental regions (Africa, Europe, Central/South Asia, Middle East Asia, East Asia, America, and Oceania).
The study begins by evaluating multiple genotype encoding schemes and several classifiers to determine the optimal inference workflow that maximises accuracy in non-admixed inference. Encoding schemes assessed include label encoding (for nucleotide sequences), one-hot encoding (for allele-specific genotypes), additive encoding (to count heterozygous allele states), and genotypic encoding (to reflect allele diversity). Three classifiers—Hidden Markov Models (HMM), Support Vector Machines (SVM), and Convolutional Neural Networks (CNN)—are evaluated given their demonstrated effectiveness in SNP genotype analyses for genetic prediction tasks.
For admixed individuals, a specialised weighted multi-label inference model is developed to estimate ancestry proportions from multiple source populations and to detect fine-scale population stratification. The admixed genotypes used in experiments are simulated across three generations using standard forward-time simulation, based on seven reference populations.
In addition, a multistage AIM identification system is proposed. This system employs a multivariate approach for variable selection, combining discriminant analysis and correlation analysis to identify informative SNPs. Metrics and methods used for identification include absolute allele frequency, linkage disequilibrium parameters (LD parameter), Cram ´e’s V, and recursive feature elimination (RFE). Metrics used for validation comprise locus informativeness, allele frequency difference distribution, and pairwise allele frequency distances.
The main assumptions adopted in this study, based exclusively on autosomal genotypes, are as follows: humans are considered a genetically homogeneous species; populations are assumed to be in Hardy-Weinberg Equilibrium (HWE), both with respect to allele frequencies and in the absence of evolutionary influences; and recombination follows the principles of Mendelian particulate inheritance.
The proposed inference approach and AIM identification system are applied to individuals surveyed in the Human Genome Diversity Panel (HGDP-CEPH), as well as to simulated admixed samples generated from HGDP reference populations. Prior to any inference tasks, principal component analysis (PCA) and F-statistics are used to validate whether the selected SNPs sufficiently represent global human diversity, ensuring that genetic relationships between individuals and populations are accurately captured and that the true population structure is not obscured or misrepresented.
Experimental results showed improved inference coverage across global intra-continental populations and high classification accuracy, particularly among genetically close groups on the Eurasian plate. This study demonstrates that leveraging genome-wide variation with machine learning enhances ancestry classification beyond what is achievable with limited AIM-based methods, reinforcing the premise that dense genomic data captures more comprehensive ancestral information. By accounting for both linkage disequilibrium (LD) and recent admixture, the proposed inference approach provides a more realistic, adaptable, and high-performing solution for admixed populations, thereby reducing the reliance on the assumption of linkage equilibrium required by current methods.
| Date of Award | 2025 |
|---|---|
| Original language | English |
| Supervisor | Dat TRAN (Supervisor) & Elisa Martinez-Marroquin (Supervisor) |
Cite this
- Standard