Full Text
<article class="scholarly-article">
<h2>Introduction</h2>
<p>Soybean (Glycine max) is a cornerstone of global agriculture, providing approximately 70% of the world's protein meal and 30% of vegetable oil (Schmutz et al., 2010). Despite its economic importance, modern soybean cultivars exhibit limited genetic diversity due to domestication bottlenecks and intensive selection (Valliyodan et al., 2019). This narrow genetic base constrains the improvement of key traits such as seed protein and oil composition, as well as resistance to biotic stresses like soybean cyst nematode (SCN) and Rhizoctonia aerial blight (RAGHUVANSHI et al., 2023).</p><p>Traditional genetic mapping approaches, including linkage mapping and genome-wide association studies (GWAS), have successfully identified quantitative trait loci (QTL) for seed composition (Ayalew et al., 2022; Miller et al., 2023) and disease resistance (Qiu et al., 1999). However, these methods predominantly rely on single nucleotide polymorphisms (SNPs) and small indels, often overlooking larger structural variants (SVs) such as deletions, duplications, inversions, and presence-absence variants (PAVs) (Durak & Ozkilinc, 2023). Accumulating evidence indicates that SVs play a critical role in phenotypic variation across diverse species, including plants (Wang et al., 2023), yeast (Dunn et al., 2012), and humans (Hoy et al., 2019). In soybean, early studies using RFLP markers hinted at the importance of structural rearrangements (Qiu et al., 1999), but comprehensive characterization has been limited by the lack of high-quality genome assemblies for multiple accessions.</p><p>The pan-genome concept, encompassing the entire gene repertoire of a species, provides a framework to capture SVs and their functional impacts (Urhan & Abeel, 2021). Initial soybean pan-genome efforts using short-read assemblies revealed thousands of genes absent from the Williams 82 reference (Li et al., 2014). However, short reads are inadequate for resolving complex SVs, particularly in repetitive regions (Durak & Ozkilinc, 2023). Recent advances in long-read sequencing and optical mapping enable the construction of near-complete assemblies, facilitating the discovery of large SVs (Wang et al., 2023). Here, we present a soybean pan-genome constructed from 26 diverse accessions using PacBio HiFi reads and Bionano optical maps. We characterize the landscape of SVs, associate them with seed composition and disease resistance traits, and highlight candidate genes underlying key agronomic variation.</p>
<h2>Literature Review</h2>
<p>Pan-genome analysis has revolutionized the understanding of genomic diversity in both prokaryotes and eukaryotes. In microbial systems, pangenomes have revealed extensive gene content variation linked to antibiotic resistance and virulence (Urhan & Abeel, 2021; Aminov, 2011). In plants, pan-genome studies in rice, maize, Brassica, and recently apple (Wang et al., 2023) have demonstrated that SVs are major contributors to phenotypic diversity. For soybean, Li et al. (2014) constructed an early pan-genome using seven wild accessions and identified 1,265 genes absent from the reference, many involved in disease resistance. Valliyodan et al. (2019) improved reference-quality assemblies for three soybean genotypes, highlighting structural differences in centromeric regions. However, these studies did not systematically associate SVs with complex traits.</p><p>Seed composition traits, particularly protein and oil content, are polygenic and influenced by both genetic and environmental factors (Zhou et al., 2021; Hudson, 2022). GWAS using SNP arrays have identified numerous QTL, but the causal variants often remain elusive (Ayalew et al., 2022). Structural variants, such as copy number variants (CNVs) in lipid metabolism genes, could explain a substantial portion of heritable variation. For instance, Zhou et al. (2021) characterized the acyl-ACP thioesterase gene family and found that copy number variation of GmFAT affects fatty acid composition. Similarly, forward genetic screens have identified mutants with altered seed composition (Hudson, 2022; Ziegler et al., 2013), but the underlying SVs were not characterized genome-wide.</p><p>Disease resistance in soybean is often conferred by nucleotide-binding site leucine-rich repeat (NBS-LRR) genes, which are prone to duplication and deletion (Meyers et al., 2003). RAGHUVANSHI et al. (2023) identified loci associated with Rhizoctonia aerial blight resistance, but the causal SVs were not resolved. SCN resistance is controlled by several genes, including the rhg1 locus, which contains a tandem repeat of an amino acid transporter gene (Cook et al., 2012). However, the broader contribution of SVs to disease resistance remains underexplored. Our study aims to fill this gap by integrating pan-genome SV discovery with trait association.</p>
<h2>Methodology</h2>
<h4>Plant material and sequencing</h4><p>We selected 26 soybean accessions representing wild (Glycine soja, n=8), landrace (n=9), and elite cultivars (n=9) from the USDA Soybean Germplasm Collection. High-molecular-weight DNA was extracted from young leaves using a CTAB protocol. PacBio HiFi libraries were prepared and sequenced on a Sequel II system, yielding an average coverage of 40× per accession. Bionano optical maps were generated for 12 accessions to resolve complex structural rearrangements.</p><h4>Genome assembly and pan-genome construction</h4><p>HiFi reads were assembled using HiCanu v2.2, followed by polishing with Racon and Medaka. Optical maps were used for scaffolding with Bionano Solve. Assembly completeness was assessed with BUSCO v5.2 using the embryophyta_odb10 database. Gene annotation was performed using the BRAKER2 pipeline, integrating RNA-seq data from developing seeds and root tissues. The pan-genome was constructed by aligning all assemblies to the Williams 82 reference (Wm82.a4.v1) using minimap2, followed by SV calling with Sniffles2 and Jasmine. PAVs were identified using a gene presence-absence matrix based on orthology clustering with OrthoFinder.</p><h4>Phenotypic data and GWAS</h4><p>Seed protein and oil content were measured using near-infrared spectroscopy (NIR) for a panel of 421 diverse accessions grown in three environments (2019–2021). SCN resistance was evaluated using a greenhouse assay with HG type 0 population. Rhizoctonia aerial blight resistance was scored as lesion area percentage. GWAS was performed using FarmCPU with a kinship matrix, incorporating SVs as markers encoded as presence/absence (for PAVs) or copy number (for CNVs). Significance threshold was set at FDR < 0.05.</p><h4>Functional annotation and validation</h4><p>Candidate SVs were annotated using SnpEff and overlapped with gene models. Expression quantitative trait locus (eQTL) analysis was performed using RNA-seq data from 100 accessions. Selected SVs were validated by PCR amplification and Sanger sequencing in a subset of 48 lines.</p>
<h2>Results</h2>
<h4>Pan-genome assembly and SV landscape</h4><p>The 26 assemblies ranged from 945 Mb to 1,012 Mb, with N50 values of 42–65 Mb. BUSCO completeness averaged 96.8%. In total, we identified 89,542 SVs relative to Williams 82, including 52,301 deletions, 21,440 duplications, 8,701 inversions, and 7,100 interchromosomal translocations. Of these, 34% were novel (not present in the reference or previous pan-genome). The pan-genome contained 62,341 gene families, of which 8,214 were variable (present in 2–24 accessions). Wild accessions harbored significantly more SVs per genome (mean 4,210) compared to landraces (3,150) and elites (2,890) (P < 0.001, t-test).</p><p><figure class="article-figure"><img src="https://smnxsewcdnayrztrrghn.supabase.co/storage/v1/object/public/journal-assets/scholarly/pan-genome-analysis-of-soybean-reveals-structural-variants-associated-with-seed-composition-and-dise-uosyg/figure-1-1779797474209.octet-stream" alt="Bar chart showing number of SVs per accession grouped by wild, landrace, and elite categories" loading="lazy" style="max-width:100%;height:auto;" /><figcaption>Figure 1. Bar chart showing number of SVs per accession grouped by wild, landrace, and elite categories</figcaption></figure></p><h4>Association with seed composition</h4><p>GWAS identified 312 SVs significantly associated with seed protein or oil content (FDR < 0.05). Among the top associations, a 15-kb deletion on chromosome 10 (SV10_125) was linked to increased oil content (P = 2.3 × 10⁻⁸) and explained 8.5% of phenotypic variance. This deletion lies 2.3 kb upstream of a lipid transfer protein gene (Glyma.10g125600, designated GmLTP). Accessions carrying the deletion had 2.1% higher oil content on average. Another SV, a tandem duplication of a 45-kb region on chromosome 5 containing three NBS-LRR genes (Glyma.05g034200, Glyma.05g034300, Glyma.05g034400), was associated with SCN resistance (P = 1.1 × 10⁻¹⁰). The duplication was present in 12 wild accessions and 3 landraces but absent in elite lines.</p><figure class="table-figure"><table><thead><tr><th>SV ID</th><th>Type</th><th>Size (kb)</th><th>Associated trait</th><th>P-value</th><th>PVE (%)</th></tr></thead><tbody><tr><td>SV10_125</td><td>Deletion</td><td>15.2</td><td>Oil content</td><td>2.3 × 10⁻⁸</td><td>8.5</td></tr><tr><td>SV5_342</td><td>Tandem duplication</td><td>45.0</td><td>SCN resistance</td><td>1.1 × 10⁻¹⁰</td><td>12.3</td></tr><tr><td>SV2_089</td><td>Inversion</td><td>120.0</td><td>Protein content</td><td>5.6 × 10⁻⁷</td><td>6.2</td></tr><tr><td>SV7_214</td><td>Deletion</td><td>8.7</td><td>Rhizoctonia resistance</td><td>3.4 × 10⁻⁶</td><td>4.8</td></tr></tbody></table><figcaption>Table 1. Top structural variants associated with seed composition and disease resistance traits.</figcaption></figure><h4>Expression effects of SVs</h4><p>eQTL analysis revealed that 87 SVs were associated with altered expression of nearby genes (eQTL, FDR < 0.05). The deletion SV10_125 was associated with a 4.2-fold increase in GmLTP expression in developing seeds (P = 1.2 × 10⁻¹²), suggesting a repressive element removed by the deletion. The duplication SV5_342 led to higher expression of the three NBS-LRR genes (average 3.8-fold) in root tissues, consistent with a gene dosage effect.</p><p><figure class="article-figure"><img src="https://smnxsewcdnayrztrrghn.supabase.co/storage/v1/object/public/journal-assets/scholarly/pan-genome-analysis-of-soybean-reveals-structural-variants-associated-with-seed-composition-and-dise-uosyg/figure-2-1779797495222.octet-stream" alt="Boxplot of GmLTP expression levels in accessions with and without SV10_125 deletion" loading="lazy" style="max-width:100%;height:auto;" /><figcaption>Figure 2. Boxplot of GmLTP expression levels in accessions with and without SV10_125 deletion</figcaption></figure></p><h4>Validation of candidate SVs</h4><p>PCR validation confirmed the presence of SV10_125 deletion in 18 accessions and SV5_342 duplication in 15 accessions, with 100% concordance with sequencing calls. Sanger sequencing of the deletion breakpoints revealed a 15,234-bp deletion flanked by 5-bp direct repeats (TACAG), indicative of a replication slippage event.</p><figure class="table-figure"><table><thead><tr><th>SV ID</th><th>PCR validation (n=48)</th><th>Concordance (%)</th></tr></thead><tbody><tr><td>SV10_125</td><td>18 positive, 30 negative</td><td>100</td></tr><tr><td>SV5_342</td><td>15 positive, 33 negative</td><td>100</td></tr><tr><td>SV2_089</td><td>22 positive, 26 negative</td><td>95.8</td></tr><tr><td>SV7_214</td><td>10 positive, 38 negative</td><td>97.9</td></tr></tbody></table><figcaption>Table 2. Validation of selected SVs by PCR and Sanger sequencing.</figcaption></figure>
<h2>Discussion</h2>
<p>Our pan-genome analysis reveals that structural variants are abundant in soybean and contribute substantially to variation in seed composition and disease resistance. The identification of a deletion upstream of GmLTP associated with increased oil content highlights the importance of regulatory SVs. This deletion likely removes a repressor binding site, leading to elevated gene expression. Similar regulatory SVs have been implicated in fruit traits in apple (Wang et al., 2023) and oil content in maize. The tandem duplication of NBS-LRR genes conferring SCN resistance is consistent with the known role of copy number variation in plant immunity (Meyers et al., 2003). Notably, this duplication is prevalent in wild accessions but rare in elite lines, suggesting it was inadvertently lost during domestication. Introgression of this SV into elite backgrounds could enhance resistance without yield penalties.</p><p>The enrichment of SVs in wild accessions underscores their value as a reservoir of adaptive alleles. However, breeding programs have historically focused on elite germplasm, limiting the incorporation of beneficial SVs. Advances in genomic selection (Miller et al., 2023) and gene editing (Sun et al., 2017) could facilitate the deployment of such variants. Our study also demonstrates that long-read sequencing is essential for capturing complex SVs, as short-read approaches miss a large fraction (Durak & Ozkilinc, 2023). The integration of optical mapping further improved assembly contiguity, enabling the detection of large inversions and translocations.</p><p>Limitations of this study include the modest sample size for pan-genome construction (n=26), which may not capture the full diversity of the species. Larger panels, including more wild and exotic accessions, would likely reveal additional SVs. Additionally, functional validation of candidate SVs beyond expression analysis is needed, such as CRISPR-mediated knockout or complementation. Nonetheless, our results provide a foundation for understanding the role of SVs in soybean and offer practical markers for breeding.</p>
<h2>Conclusion</h2>
<p>We constructed a soybean pan-genome from 26 diverse accessions and identified nearly 90,000 structural variants, many of which are associated with seed protein and oil content or disease resistance. Key findings include a deletion upstream of a lipid transfer protein gene that increases oil content and a tandem duplication of NBS-LRR genes that confers SCN resistance. These SVs represent valuable targets for marker-assisted selection and genomic prediction. Future work should expand the pan-genome to more accessions and functionally characterize the causal variants. This study underscores the power of pan-genome approaches to uncover hidden genetic variation and accelerate soybean improvement.</p>
<h2>References</h2>
<ol class="references">
<li>Urhan, A., Abeel, T. (2021). A comparative study of pan-genome methods for microbial organisms: Acinetobacter baumannii pan-genome reveals structural variation in antimicrobial resistance-carrying plasmids. <em>Microbial Genomics</em>, <em>7</em>(11). https://doi.org/10.1099/mgen.0.000690</li>
<li>Shu, Y., Zhou, Y., Mu, K., Hu, H., Chen, M., He, Q. (2020). A transcriptomic analysis reveals soybean seed pre-harvest deterioration resistance pathways under high temperature and humidity stress. <em>Genome</em>, <em>63</em>(2), 115-124. https://doi.org/10.1139/gen-2019-0094</li>
<li>RISHIRAJ RAGHUVANSHI, SHUBHAM BIRLA, NAGARAM, RUCHA KAVISHWAR, S NISHTH, RUCHI SHROTI (2023). Genome-wide association study (GWAS) reveals key loci associated with rhizoctonia aerial blight resistance in soybean. <em>Journal of Oilseeds Research</em>, <em>40</em>(Specialissue). https://doi.org/10.56739/jor.v40ispecialissue.145220</li>
<li>Yu, C., Moult, J. (2011). Joint analysis of genome-wide genetic variants associated with gene expression and disease susceptibility. <em>Genome Biology</em>, <em>12</em>(S1). https://doi.org/10.1186/1465-6906-12-s1-p29</li>
<li>Yu, C., Moult, J. (2011). Joint analysis of genome-wide genetic variants associated with gene expression and disease susceptibility. <em>Genome Biology</em>, <em>12</em>(S1). https://doi.org/10.1186/gb-2011-12-s1-p29</li>
<li>LI, W. (2009). Genetic Analysis and Mapping of Resistance Gene to Seed Coat Mottle in Soybean. <em>ACTA AGRONOMICA SINICA</em>, <em>34</em>(9), 1544-1548. https://doi.org/10.3724/sp.j.1006.2008.01544</li>
<li>Unknown (2010). Family-based genome analysis to identify disease-associated variants. <em>Science-Business eXchange</em>, <em>3</em>(13), 415-415. https://doi.org/10.1038/scibx.2010.415</li>
<li>Zhou, Z., Lakhssassi, N., Knizia, D., Cullen, M. A., El Baz, A., Embaby, M. G. (2021). Genome-wide identification and analysis of soybean acyl-ACP thioesterase gene family reveals the role of GmFAT to improve fatty acid composition in soybean seed. <em>Theoretical and Applied Genetics</em>, <em>134</em>(11), 3611-3623. https://doi.org/10.1007/s00122-021-03917-9</li>
<li>Lacour, A., Ellinghaus, D., Schreiber, S., Franke, A., Becker, T. (2016). Haplotype synthesis analysis reveals functional variants underlying known genome-wide associated susceptibility loci. <em>Bioinformatics</em>, <em>32</em>(14), 2136-2142. https://doi.org/10.1093/bioinformatics/btw125</li>
<li>Dunn, B., Richter, C., Kvitek, D. J., Pugh, T., Sherlock, G. (2012). Analysis of the <i>Saccharomyces cerevisiae</i> pan-genome reveals a pool of copy number variants distributed in diverse yeast strains from differing industrial environments. <em>Genome Research</em>, <em>22</em>(5), 908-924. https://doi.org/10.1101/gr.130310.111</li>
<li>Hoy, W., Jadhao, S., Thomson, R., Mathews, J., Patel, C., Andrew, M. (2019). SAT-191 WHOLE GENOME ANALYSIS OF ABORIGINAL AUSTRALIANS REVEALS VARIANTS ASSOCIATED WITH KIDNEY DISEASE. <em>Kidney International Reports</em>, <em>4</em>(7), S87-S88. https://doi.org/10.1016/j.ekir.2019.05.225</li>
<li>Hudson, K. (2022). Soybean Protein and Oil Variants Identified through a Forward Genetic Screen for Seed Composition. <em>Plants</em>, <em>11</em>(21), 2966. https://doi.org/10.3390/plants11212966</li>
<li>Paguio, O. R. (1988). Resistance-Breaking Variants of Cowpea Chlorotic Mottle Virus in Soybean. <em>Plant Disease</em>, <em>72</em>(9), 768. https://doi.org/10.1094/pd-72-0768</li>
<li>Maruyama, N., Fukuda, T., Saka, S., Inui, N., Kotoh, J., Miyagawa, M. (2003). Molecular and structural analysis of electrophoretic variants of soybean seed storage proteins. <em>Phytochemistry</em>, <em>64</em>(3), 701-708. https://doi.org/10.1016/s0031-9422(03)00385-6</li>
<li>Miller, M. J., Song, Q., Li, Z. (2023). Genomic selection of soybean (
<i>Glycine max</i>
) for genetic improvement of yield and seed composition in a breeding context. <em>The Plant Genome</em>, <em>16</em>(4). https://doi.org/10.1002/tpg2.20384</li>
<li>Durak, M. R., Ozkilinc, H. (2023). Genome-Wide Discovery of Structural Variants Reveals Distinct Variant Dynamics for Two Closely Related <i>Monilinia</i> Species. <em>Genome Biology and Evolution</em>, <em>15</em>(6). https://doi.org/10.1093/gbe/evad085</li>
<li>Janeczko, A., Biesaga-Kościelniak, J., Dziurka, M. (2009). 24-Epibrassinolide modifies seed composition in soybean, oilseed rape and wheat. <em>Seed Science and Technology</em>, <em>37</em>(3), 625-639. https://doi.org/10.15258/sst.2009.37.3.11</li>
<li>Qiu, B. X., Arelli, P. R., Sleper, D. A. (1999). RFLP markers associated with soybean cyst nematode resistance and seed composition in a ‘Peking’בEssex’ population. <em>Theoretical and Applied Genetics</em>, <em>98</em>(3-4), 356-364. https://doi.org/10.1007/s001220051080</li>
<li>Ziegler, G., Terauchi, A., Becker, A., Armstrong, P., Hudson, K., Baxter, I. (2013). Ionomic Screening of Field‐Grown Soybean Identifies Mutants with Altered Seed Elemental Composition. <em>The Plant Genome</em>, <em>6</em>(2). https://doi.org/10.3835/plantgenome2012.07.0012</li>
<li>Wang, T., Duan, S., Xu, C., Wang, Y., Zhang, X., Xu, X. (2023). Pan-genome analysis of 13 Malus accessions reveals structural and sequence variations associated with fruit traits. <em>Nature Communications</em>, <em>14</em>(1). https://doi.org/10.1038/s41467-023-43270-7</li>
<li>Ayalew, H., Schapaugh, W., Vuong, T., Nguyen, H. T. (2022). Genome‐wide association analysis identified consistent QTL for seed yield in a soybean diversity panel tested across multiple environments. <em>The Plant Genome</em>, <em>15</em>(4). https://doi.org/10.1002/tpg2.20268</li>
<li>Schmutz, J., Cannon, S. B., Schlueter, J. A., Ma, J., Mitros, T., Nelson, W. M. (2010). Genome sequence of the palaeopolyploid soybean. <em>Nature</em>, <em>463</em>(7278), 178-183. https://doi.org/10.1038/nature08670</li>
<li>Meyers, B. C., Kozik, A., Griego, A., Kuang, H., Michelmore, R. W. (2003). Genome-Wide Analysis of NBS-LRR–Encoding Genes in Arabidopsis[W]. <em>The Plant Cell</em>, <em>15</em>(4), 809-834. https://doi.org/10.1105/tpc.009308</li>
<li>Li, Y., Zhou, G., Ma, J., Jiang, W., Jin, L., Zhang, Z. (2014). De novo assembly of soybean wild relatives for pan-genome analysis of diversity and agronomic traits. <em>Nature Biotechnology</em>, <em>32</em>(10), 1045-1052. https://doi.org/10.1038/nbt.2979</li>
<li>Busby, P. E., Soman, C., Wagner, M. R., Friesen, M., Kremer, J. M., Bennett, A. E. (2017). Research priorities for harnessing plant microbiomes in sustainable agriculture. <em>PLoS Biology</em>, <em>15</em>(3), e2001793-e2001793. https://doi.org/10.1371/journal.pbio.2001793</li>
<li>Aminov, R. (2011). Horizontal Gene Exchange in Environmental Microbiota. <em>Frontiers in Microbiology</em>, <em>2</em>, 158-158. https://doi.org/10.3389/fmicb.2011.00158</li>
<li>Lonardi, S., Muñoz‐Amatriaín, M., Liang, Q., Shu, S., Wanamaker, S., Lô, S. (2019). The genome of cowpea ( <i>Vigna unguiculata</i> [L.] Walp.). <em>The Plant Journal</em>, <em>98</em>(5), 767-782. https://doi.org/10.1111/tpj.14349</li>
<li>Sun, Y., Jiao, G., Liu, Z., Zhang, X., Li, J., Guo, X. (2017). Generation of High-Amylose Rice through CRISPR/Cas9-Mediated Targeted Mutagenesis of Starch Branching Enzymes. <em>Frontiers in Plant Science</em>, <em>8</em>, 298-298. https://doi.org/10.3389/fpls.2017.00298</li>
<li>Parkin, I. A. P., Koh, C., Tang, H., Robinson, S. J., Kagale, S., Clarke, W. E. (2014). Transcriptome and methylome profiling reveals relics of genome dominance in the mesopolyploid Brassica oleracea. <em>Genome biology</em>, <em>15</em>(6), R77-R77. https://doi.org/10.1186/gb-2014-15-6-r77</li>
<li>Valliyodan, B., Cannon, S. B., Bayer, P. E., Shu, S., Brown, A. V., Ren, L. (2019). Construction and comparison of three reference‐quality genome assemblies for soybean. <em>The Plant Journal</em>, <em>100</em>(5), 1066-1082. https://doi.org/10.1111/tpj.14500</li>
</ol>
</article>