Abstract
Improving spring wheat yield is essential for global food security, yet breeding is constrained by biological complexity and statistical modeling limitations. Biologically, yield lacks a deterministic gene; selection is inefficient due to complex source-sink trade-offs and genotype-by-environment interactions. Statistically, conventional genomic selection (GS) models (like GBLUP) assume loci possess small, uniform effects, failing to leverage biological priors. High-density genotyping partially compensates but imposes a severe economic burden. Resolving biological mechanisms provides the foundation to optimize models, enabling the integration of major-effect markers for cost reduction. This dissertation establishes a biologically-informed GS framework across three phases.To demystify yield formation, the first phase analyzed FLOWERING LOCUS T (FT) homoeologs, which sit at integrative nodes of signaling pathways and exhibit pleiotropy across yield components. Using near-isogenic lines, a diversity panel, and CRISPR-Cas9 mutagenesis, the physiological trade-offs of FT-B1 and FT-D1 were elucidated. Mutant alleles at both loci expanded sink capacity but induced a compensatory reduction in grain weight. Crucially, only FT-D1 translated this expansion into a significant yield advantage by driving a synergistic taller and later phenotype, extending the resource acquisition window to fulfill the expanded sink.
Having identified these genes as decisive priors, the second phase integrated diagnostic markers for developmental genes (FT, Ppd, Rht, Vrn) as fixed effects into Reproducing Kernel Hilbert Spaces (RKHS) models. Injecting this knowledge explicitly captured large-effect variance. This approach enhanced predictive abilities, yielding a 13.6% increase in grain yield prediction accuracy over the baseline model and demonstrating that trait-specific marker combinations outperform single-gene strategies.
Because these biological factors captured substantial predictive variance, reliance on high-density background markers was reduced. The third phase demonstrated that a mid-density 4K Genotyping-by-Sequencing platform effectively replaced the costly 90K SNP array. The 4K SNP-only model achieved predictive abilities comparable to the 90K array due to marker-density saturation. Integrating diagnostic markers into the 4K panel further improved predictions for heading date and plant height, with cross-population validations confirming robust transferability.
By translating gene function into mechanism-guided models, this research tries to deliver an actionable, high-precision, and low-cost genomic breeding solution.