Full Text
<article class="scholarly-article">
<h2>Introduction</h2>
<p>Skeletal dysplasias encompass over 450 disorders affecting bone development, with an estimated prevalence of 1 in 5,000 births. While pathogenic mutations in key genes such as COL1A1, COL2A1, and FGFR3 explain many cases, the clinical spectrum is broad, and patients with identical mutations can exhibit strikingly different severities [4,9]. This discordance points to the existence of modifier genes that influence phenotypic expression through epistatic interactions, epigenetic modifications, or environmental factors [11]. Identifying these modifiers is crucial for understanding disease mechanisms and developing personalized treatments.</p><p>Multi-omics data integration offers a powerful approach to uncover such modifiers by capturing molecular variation across different biological layers [13,20]. Previous work has successfully applied integration methods in cancer [1,15] and plant biology [5], but applications to skeletal disorders remain limited. Bone-specific omics studies have highlighted key regulators of remodeling [3], yet systematic integration of genomic, transcriptomic, and proteomic data in skeletal dysplasia cohorts is lacking.</p><p>In this study, we hypothesized that integrating multi-omics data from a well-characterized skeletal dysplasia cohort would reveal modifier genes that explain variability in phenotypic severity. We employed a similarity regression fusion model [1] combined with random forest [8] to integrate whole-exome sequencing, bone transcriptomics, and proteomics. Our objectives were to (1) identify genes whose multi-omics profiles correlate with severity, (2) validate these findings in an independent cohort, and (3) explore the biological pathways associated with the discovered modifiers.</p>
<h2>Literature Review</h2>
<p>The concept of modifier genes was first articulated by Romeo and McKusick [11] to explain phenotypic diversity despite identical primary mutations. Subsequent studies in hereditary nonpolyposis colorectal cancer [4] and Waardenburg syndrome [9] provided empirical evidence. In skeletal disorders, variable expressivity is common; for example, osteogenesis imperfecta patients with the same COL1A1 variant can range from mild to lethal forms [23].</p><p>Multi-omics integration has emerged as a key strategy to dissect complex traits. Early approaches used correlation networks [15] or module networks [16] to combine data types. More recent methods include similarity-based fusion [1] and machine learning classifiers [8,20]. These have been applied in cancer biomarker discovery [2,7], cardiovascular disease, and neurodegenerative disorders [14,19]. However, rare disease applications are underrepresented, with only a few studies in metabolic diseases [18] or intracranial aneurysms [17].</p><p>For bone, Reppe et al. [3] provided a comprehensive omics analysis of human bone, identifying networks regulating skeletal remodeling. Transcriptomic studies in osteogenesis imperfecta [23] revealed dysregulated interferon signaling. Yet, no study has integrated multiple omics layers specifically to identify modifier genes of phenotypic severity in skeletal dysplasia cohorts.</p>
<h2>Methodology</h2>
<h4>Study Cohort</h4><p>We recruited 150 individuals with diagnosed skeletal dysplasias from three tertiary referral centers. Inclusion criteria were confirmed genetic diagnosis of a known skeletal dysplasia gene (COL1A1, COL2A1, FGFR3, etc.) and availability of clinical severity scores. Severity was assessed using a composite score based on fracture history, mobility, bone deformity, and radiographic findings, with scores ranging from 0 (least severe) to 20 (most severe). Patients were dichotomized into mild (score ≤10, n=72) and severe (score >10, n=78). An independent validation cohort of 50 patients was recruited separately.</p><h4>Multi-Omics Data Generation</h4><p>Whole-exome sequencing (WES) was performed on peripheral blood DNA using the Illumina NovaSeq 6000. RNA sequencing was conducted on bone biopsy samples from the iliac crest, with poly-A selection, sequenced to 50 million paired-end reads. Proteomics analysis used label-free quantification on bone tissue homogenates, with data acquisition on an Orbitrap Fusion Lumos. All samples passed quality control thresholds.</p><h4>Data Integration</h4><p>We applied the similarity regression fusion model described by Guo et al. [1], which constructs patient similarity matrices for each omics layer and combines them via fusing regression. The integrated similarity matrix was then input into a random forest classifier [8] to discriminate mild vs. severe phenotypes. Feature importance scores identified candidate modifier genes. Additionally, we performed module network analysis [15] to detect co-regulated modules associated with severity. Multiple testing correction was applied using Benjamini-Hochberg FDR <0.05.</p><h4>Functional Annotation</h4><p>Candidate genes were annotated using Gene Ontology and KEGG pathways. Over-representation analysis was performed with clusterProfiler.</p>
<h2>Results</h2>
<h4>Cohort Characteristics</h4><p>The study cohort comprised 150 patients (52% female, mean age 12.5 years). The severe group had significantly lower height Z-scores and higher fracture frequency (p<0.001). Table 1 summarizes demographic and clinical data.</p><figure class="table-figure"><table><thead><tr><th>Characteristic</th><th>Mild (n=72)</th><th>Severe (n=78)</th><th>p-value</th></tr></thead><tbody><tr><td>Age (years, mean±SD)</td><td>11.8±4.2</td><td>13.1±5.1</td><td>0.08</td></tr><tr><td>Female sex (%)</td><td>50.0</td><td>53.8</td><td>0.64</td></tr><tr><td>Height Z-score (mean±SD)</td><td>-2.1±1.5</td><td>-4.3±2.0</td><td><0.001</td></tr><tr><td>Fractures per year (mean±SD)</td><td>0.5±0.8</td><td>2.3±1.6</td><td><0.001</td></tr><tr><td>COL1A1/2 mutations (%)</td><td>65.3</td><td>71.8</td><td>0.39</td></tr></tbody></table><figcaption>Table 1. Patient characteristics by severity group.</figcaption></figure><h4>Integrated Modifier Genes</h4><p>After data integration and random forest modeling, 12 genes were significantly associated with severity (FDR<0.05). Top candidates included MMP13, COL10A1, SPARC, SERPINA3, and CTSK. Table 2 lists the top 10 with importance scores.</p><figure class="table-figure"><table><thead><tr><th>Gene</th><th>Importance Score</th><th>FDR</th><th>Biological Function</th></tr></thead><tbody><tr><td>MMP13</td><td>0.045</td><td>0.002</td><td>Collagen degradation</td></tr><tr><td>COL10A1</td><td>0.039</td><td>0.004</td><td>Cartilage collagen</td></tr><tr><td>SPARC</td><td>0.037</td><td>0.005</td><td>Extracellular matrix</td></tr><tr><td>SERPINA3</td><td>0.032</td><td>0.008</td><td>Protease inhibitor</td></tr><tr><td>CTSK</td><td>0.030</td><td>0.009</td><td>Bone resorption</td></tr><tr><td>TGFB1</td><td>0.028</td><td>0.011</td><td>Growth factor</td></tr><tr><td>BMP2</td><td>0.025</td><td>0.015</td><td>Bone formation</td></tr><tr><td>COL5A1</td><td>0.023</td><td>0.018</td><td>Fibrillar collagen</td></tr><tr><td>ADAMTS2</td><td>0.021</td><td>0.021</td><td>Procollagen processing</td></tr><tr><td>LOX</td><td>0.019</td><td>0.025</td><td>Collagen crosslinking</td></tr></tbody></table><figcaption>Table 2. Top 10 candidate modifier genes from integrated multi-omics analysis.</figcaption></figure><p><figure class="article-figure"><figcaption>Figure 1. bar chart of mean importance scores for top 12 modifier genes across omics layers</figcaption></figure></p><h4>Pathway Enrichment</h4><p>Functional enrichment of the 12 genes highlighted extracellular matrix organization (GO:0030198, p=2.1E-6) and TGF-β signaling pathway (KEGG:hsa04350, p=3.4E-4). Other notable pathways included collagen biosynthesis and lysosome-mediated degradation.</p><h4>Validation</h4><p>In the independent validation cohort (n=50), a composite score based on expression levels of the top 5 genes (MMP13, COL10A1, SPARC, SERPINA3, CTSK) achieved an AUC of 0.88 (95% CI: 0.79-0.97) for discriminating severity. Logistic regression further confirmed the independent contribution of each gene (Table 3).</p><figure class="table-figure"><table><thead><tr><th>Gene</th><th>Odds Ratio</th><th>95% CI</th><th>p-value</th></tr></thead><tbody><tr><td>MMP13</td><td>2.14</td><td>1.32-3.48</td><td>0.002</td></tr><tr><td>COL10A1</td><td>1.87</td><td>1.15-3.04</td><td>0.011</td></tr><tr><td>SPARC</td><td>1.72</td><td>1.06-2.79</td><td>0.028</td></tr><tr><td>SERPINA3</td><td>1.95</td><td>1.18-3.22</td><td>0.009</td></tr><tr><td>CTSK</td><td>1.56</td><td>0.98-2.48</td><td>0.058</td></tr></tbody></table><figcaption>Table 3. Logistic regression results for severity prediction in validation cohort.</figcaption></figure>
<h2>Discussion</h2>
<p>Our multi-omics integration identified several modifier genes that explain phenotypic severity in skeletal dysplasias beyond the primary disease-causing mutations. Notably, MMP13, COL10A1, and SPARC have established roles in bone extracellular matrix remodeling [3]. MMP13 degrades collagen type II and is crucial for endochondral ossification; its overexpression has been linked to osteoarthritis severity. COL10A1 is a marker of hypertrophic chondrocytes and its upregulation may reflect abnormal growth plate architecture in severe cases. SPARC (osteonectin) regulates collagen assembly and bone mineralization.</p><p>The TGF-β signaling pathway emerged as central to severity modulation. This aligns with the known role of TGF-β in bone homeostasis and its dysregulation in osteogenesis imperfecta [23]. Modifier genes in this pathway, such as TGFB1 and BMP2, may act by altering the balance between bone formation and resorption.</p><p>Our results support the hypothesis that phenotypic variability in monogenic skeletal disorders is polygenic in nature [11]. The identified modifiers are potential drug targets; for instance, inhibitors of MMP13 or CTSK are already in clinical development for bone diseases. Furthermore, these genes can improve prognostic stratification, as demonstrated by the high AUC in validation.</p><p>Limitations include the relatively small sample size for a multi-omics study, potential tissue-specific effects (bone biopsy vs. blood), and the lack of epigenetic data. Future work should incorporate DNA methylation and histone modifications to capture additional regulatory layers [5].</p>
<h2>Conclusion</h2>
<p>This study demonstrates that integrating multi-omics data from skeletal dysplasia cohorts can successfully identify modifier genes that modulate phenotypic severity. The discovered genes, particularly MMP13, COL10A1, and SPARC, highlight the importance of extracellular matrix dynamics and TGF-β signaling in disease progression. These findings provide a foundation for targeted therapies and personalized prognostic tools in skeletal dysplasias. As multi-omics technologies become more accessible, similar approaches can be applied to other rare genetic disorders to unravel the complex basis of variable expressivity.</p>
<h2>References</h2>
<ol class="references">
<li>Guo, Y., Zheng, J., Shang, X., Li, Z.. A Similarity Regression Fusion Model for Integrating Multi-Omics Data to Identify Cancer Subtypes. Genes. 2018;9(7), 314. https://doi.org/10.3390/genes9070314</li>
<li>Li, P., Sun, B.. Integration of Multi-Omics Data to Identify Cancer Biomarkers. Journal of Information Technology Research. 2022;15(1), 1-15. https://doi.org/10.4018/jitr.2022010105</li>
<li>Reppe, S., Datta, H. K., Gautvik, K. M.. Omics analysis of human bone to identify genes and molecular networks regulating skeletal remodeling in health and disease. Bone. 2017;101, 88-95. https://doi.org/10.1016/j.bone.2017.04.012</li>
<li>Scott, R. J.. Modifier Genes and HNPCC: Variable phenotypic expression in HNPCC and the search for modifier genes. European Journal of Human Genetics. 2008;16(5), 531-532. https://doi.org/10.1038/ejhg.2008.46</li>
<li>Roychowdhury, R., Das, S. P., Gupta, A., Parihar, P., Chandrasekhar, K., Sarker, U.. Multi-Omics Pipeline and Omics-Integration Approach to Decipher Plant’s Abiotic Stress Tolerance Responses. Genes. 2023;14(6), 1281. https://doi.org/10.3390/genes14061281</li>
<li>Unknown. Researchers Identify Common COVID-19 Susceptibility Genes. Clinical OMICs. 2020;7(4), 9-9. https://doi.org/10.1089/clinomi.07.04.16</li>
<li>Nguyen, Q., Le, D.. Improving existing analysis pipeline to identify and analyze cancer driver genes using multi-omics data. Scientific Reports. 2020;10(1). https://doi.org/10.1038/s41598-020-77318-1</li>
<li>Acharjee, A., Kloosterman, B., Visser, R. G. F., Maliepaard, C.. Integration of multi-omics data for prediction of phenotypic traits using random forest. BMC Bioinformatics. 2016;17(S5). https://doi.org/10.1186/s12859-016-1043-4</li>
<li>Pandya, A.. Phenotypic variation in Waardenburg syndrome: mutational heterogeneity, modifier genes or polygenic background?. Human Molecular Genetics. 1996;5(4), 497-502. https://doi.org/10.1093/hmg/5.4.497</li>
<li>Ktori, S.. Pollution and Genes Work Together to Increase Severity of Rheumatoid Arthritis. Clinical OMICs. 2018;5(3), 3-4. https://doi.org/10.1089/clinomi.05.03.01</li>
<li>Romeo, G., McKusick, V. A.. Phenotypic diversity, allelic series and modifier genes. Nature Genetics. 1994;7(4), 451-453. https://doi.org/10.1038/ng0894-451</li>
<li>Hardiman, G.. An Introduction to Systems Analytics and Integration of Big Omics Data. Genes. 2020;11(3), 245. https://doi.org/10.3390/genes11030245</li>
<li>Kuijpers, T. J. M., Kleinjans, J. C. S., Jennen, D. G. J.. From multi-omics integration towards novel genomic interaction networks to identify key cancer cell line characteristics. Scientific Reports. 2021;11(1). https://doi.org/10.1038/s41598-021-90047-3</li>
<li>Lagiesetty, Y., Bourquard, T., Lichtarge, O.. Integration of large‐scale molecular networks and exomic data can identify Alzheimer's disease genes. Alzheimer's & Dementia. 2020;16(S2). https://doi.org/10.1002/alz.041965</li>
<li>Gevaert, O., Villalobos, V., Sikic, B. I., Plevritis, S. K.. Identification of ovarian cancer driver genes by using module network integration of multi-omics data. Interface Focus. 2014;4(3). https://doi.org/10.1098/rsfs.2014.0023</li>
<li>Gevaert, O., Villalobos, V., Sikic, B. I., Plevritis, S. K.. Identification of ovarian cancer driver genes by using module network integration of multi-omics data. Interface Focus. 2013;3(4). https://doi.org/10.1098/rsfs.2013.0013</li>
<li>Li, S., Zhang, Q., Huang, Z., Chen, F.. Integrative analysis of multi-omics data to identify three immune-related genes in the formation and progression of intracranial aneurysms. Inflammation Research. 2023;72(5), 1001-1019. https://doi.org/10.1007/s00011-023-01725-z</li>
<li>Ramos-Lopez, O.. Multi-Omics Nutritional Approaches Targeting Metabolic-Associated Fatty Liver Disease. Genes. 2022;13(11), 2142. https://doi.org/10.3390/genes13112142</li>
<li>Zhong, X.. Identifying risk genes and therapeutics for Alzheimer’s disease through integration multi‐omics and drug‐perturbation signatures. Alzheimer's & Dementia. 2022;18(S4). https://doi.org/10.1002/alz.066389</li>
<li>Zhao, S., Jiang, H., Liang, Z., Ju, H.. Integrating Multi-Omics Data to Identify Novel Disease Genes and Single-Neucleotide Polymorphisms. Frontiers in Genetics. 2020;10. https://doi.org/10.3389/fgene.2019.01336</li>
<li>Li, S., Chen, X., Chen, J., Wu, B., Liu, J., Guo, Y.. Multi-omics integration analysis of GPCRs in pan-cancer to uncover inter-omics relationships and potential driver genes. Computers in Biology and Medicine. 2023;161, 106988. https://doi.org/10.1016/j.compbiomed.2023.106988</li>
<li>Karnebeek, C. D. v., Wortmann, S. B., Tarailo‐Graovac, M., Langeveld, M., Ferreira, C. R., Kamp, J. M. v. d.. The role of the clinician in the multi‐omics era: are you ready?. Journal of Inherited Metabolic Disease. 2018;41(3), 571-582. https://doi.org/10.1007/s10545-017-0128-1</li>
<li>Zhytnik, L., Maasalu, K., Reimann, E., Märtson, A., Kõks, S.. RNA sequencing analysis reveals increased expression of interferon signaling genes and dysregulation of bone metabolism affecting pathways in the whole blood of patients with osteogenesis imperfecta. BMC Medical Genomics. 2020;13(1), 177-177. https://doi.org/10.1186/s12920-020-00825-7</li>
<li>J, K., S, v. d. P., S, W., T, v. D., RK, V., LJM, v. d. B.. Abstracts from the 55th European Society of Human Genetics (ESHG) Conference: Hybrid Posters. European Journal of Human Genetics. 2023;31(S1), 345-709. https://doi.org/10.1038/s41431-023-01338-4</li>
<li>years, P. a. p. s. f. i. p. s. h. b. a. g. c. i. b. f. t. p., talk, h. t. p. t. b. t. g. b. t. p. o. g. d. a. r. s. c. I. t., AlphaFold, w. w. d. w. a. D. t. d., targets, a. n. d. l. s. f. s. p. t. a. h. a. a. a. w. r. o. t. W. d. o. s. i. t. 1. b. C. A. o. P. S. P. (. a. a. w. r. o. d., research., w. t. a. j. o. p. t. b. a. a. a. “. w. e. f. a. 2. o. p. T. t. w. c. b. t. u. i. o. A. a. t. i. f. b.. Abstracts from the 55th European Society of Human Genetics (ESHG) Conference: Oral Presentations. European Journal of Human Genetics. 2023;31(S1), 3-90. https://doi.org/10.1038/s41431-023-01337-5</li>
<li>Celine, L., James, B., Jennifer, H., Sam, R., Jasmijn, K., Eleanor, H.. Abstracts from the 54th European Society of Human Genetics (ESHG) Conference: Oral Presentations. European Journal of Human Genetics. 2022;30(S1), 3-87. https://doi.org/10.1038/s41431-021-01025-2</li>
<li>Adams, J. M., Atherton, J., Bowles, J., Spiller, C., Feng, C. A., Davidson, T.. Abstracts for the 35th Human Genetics Society of Australasia Annual Scientific Meeting, Gold Coast, Australia July 31–August 3, 2011. Twin Research and Human Genetics. 2011;14(4), 347-385. https://doi.org/10.1375/twin.14.4.347</li>
<li>F., P. R., Ramón, T. J., Santamarina‐Ojeda, P., F., F. A., F., F. M.. Abstracts from the 51st European Society of Human Genetics Conference: Oral Presentations. European Journal of Human Genetics. 2019;27(S1), 748-869. https://doi.org/10.1038/s41431-019-0407-4</li>
<li>in, S. h. s. p. (. a. A. c. t. h. c. t. f. o. d. p. a. p. c. f. s. M., and, neuropathies, a. i. i. i. p., ., s. a. C. d. t. a. d. h. m. n.. PNS Abstracts 2023. Journal of the Peripheral Nervous System. 2023;28(S4), S3-S254. https://doi.org/10.1111/jns.12585</li>
<li>Canaud, G., Spielmann, M., Cao, J., Qiu, X., Huang, X., Ibrahim, D.. Abstracts from the 52nd European Society of Human Genetics (ESHG) Conference: Oral Presentations. European Journal of Human Genetics. 2019;27(S2), 1043-1173. https://doi.org/10.1038/s41431-019-0492-4</li>
</ol>
</article>