Abstract
Multi-omics profiling technologies, measuring genomics, epigenomics, transcriptomics, proteomics, enable characterization of complex biological systems across multiple molecular layers. The large-scale multi-omics datasets provide complementary views of complex molecular mechanisms of complex diseases, yet integrative and interpretable analysis of the large-scale datasets to identify key disease-associated biomarkers and dysfunctional signaling pathways remains an open problem due to the large and complex molecular signaling systems. To resolve the fundamental challenges in integrative and interpretable multi-omics data analysis, in this dissertation, I developed several Graph-AI frameworks that systematically integrate multi-omics data with knowledge graphs, and analyzed the data with novel graph-based learning and reasoning models. The central idea is to represent large-scale multi-omics data with comprehensive textual biomedical prior knowledge and signaling networks as structured graphs that explicitly model relationships among molecular entities, cellular processes, exposures, phenotypes and diseases, which can be analyzed by novel graph AI models. Specifically, this dissertation firstly addresses two fundamental challenges: integrating heterogeneous multi-omics measurements while preserving biologically meaningful signaling relationships, and harmonizing fragmented biomedical knowledge scattered across diverse databases with inconsistent nomenclature systems. Conventional multi-omics analysis pipelines typically produce tabular feature matrices that do not capture regulatory dependencies among molecular entities, making it difficult to directly apply Graph-AI models for mechanistic interpretation. To address this limitation, this dissertation develops mosGraphGen, a multi-omics signaling graph generation pipeline that systematically maps genomics, epigenomics, transcriptomics, and proteomics data onto biologically meaningful multi-level signaling networks spanning promoters, genes, transcripts, proteins, and pathways. mosGraphGen enables the transformation of heterogeneous molecular measurements into graph-structured representations that preserve regulatory relationships across molecular layers and can be directly used for Graph-AI modeling. In scientific discovery, prior knowledge is also essential for interpret omics data analysis results. Herein, we aim to develop graph AI models combining omics data with comprehensive and consistent biomedical prior knowledge. However, existing knowledge resources are fragmented across heterogeneous databases with partially overlapping identifiers and varying entity definitions, limiting interoperability and biological resolution. To overcome these limitations, this dissertation develops BioMedGraphica, an all-in-one biomedical knowledge graph platform that harmonizes fragmented resources by integrating diverse entity types and relationship types spanning molecular biology, environmental exposures, phenotypes, diseases and drugs. BioMedGraphica further introduces the text-numeric graph (TNG) representation, which aligns structured textual prior knowledge with quantitative biomedical features to support graph-based learning and reasoning. By linking multi-omics measurements with harmonized biomedical knowledge across multiple biological scales, mosGraphGen and BioMedGraphica together establish a reusable graph-based data foundation that supports downstream Graph-AI modeling, multimodal learning, and integrative biological discovery. In parallel with the development of methods for multi-omics data integration and representation, this dissertation develops a series of Graph-AI models that advance integrative multi-omics data analysis across multiple biological scales. For tissue-level multi-omics analysis, M3NetFlow introduces a multi-scale, multi-hop, multi-omics graph neural network framework that captures hierarchical signaling structures through pathway-level subgraphs and multi-hop molecular interactions, enabling both hypothesis-guided and data-driven identification of disease-associated targets and pathways. Building upon this formulation, mosGraphFlow incorporates multi-omics signaling graphs (mosGraphs) that explicitly represent regulatory dependencies across molecular layers, enabling improved modeling of hierarchical signaling relationships and enhancing interpretability of disease-relevant signaling mechanisms. To further explore biological feature representation, this dissertation develops multimodal graph-language approaches that integrate sequence-based and textual biological knowledge with graph learning. GraphSeqLM enhances graph neural networks by incorporating DNA, RNA, and protein sequence embeddings generated from large language models, enabling joint modeling of sequence-derived biological properties and signaling network topology for improved predictive performance in bulk multi-omics studies. Complementing this direction, GALAX integrates graph neural networks into language model reasoning through reinforcement learning guided by a graph process reward model, enabling interpretable subgraph reasoning that combines quantitative multi-omic evidence, structured biological knowledge, and language-based inference to support transparent hypothesis generation and biologically grounded target discovery. Extending multi-modal (text + omics) graph learning to single-cell omics data, this dissertation further introduces Text-Omic Signaling Graphs (TOSGs) and constructs OmniCellTOSG, a large-scale resource derived from millions of single-cell and single-nucleus transcriptomic profiles across tissues and disease contexts. Based on this representation, CellTOSG-FM is developed as a multimodal Graph-Language Foundation Model (GLFM) that jointly learns from textual biological knowledge, quantitative omic data, and signaling network topology, enabling improved cellular representation learning, phenotype prediction, and signaling pathway inference. Together, these Graph-AI models demonstrate how novel graph AI modeling and reasoning can be integrated to support interpretable and scalable multi-omics data analysis across biological levels. This dissertation establishes Graph-AI frameworks that unify multi-omics data integration, biomedical knowledge graph construction, graph representation learning, multimodal foundation modeling, and interpretable reasoning within a coherent analytical paradigm. These contributions advance the development of interpretable and biologically grounded artificial intelligence methods for precision medicine and provide a generalizable foundation for discovering molecular mechanisms, biomarkers, and therapeutic targets from complex biomedical data.
Committee Chair
Fuhai Li
Committee Members
Jin Zhang; Lei Liu; Michael Province; Yixin Chen
Degree
Doctor of Philosophy (PhD)
Author's Department
Biology and Biomedical Sciences
Document Type
Dissertation
Date of Award
6-24-2026
Language
English (en)
DOI
https://doi.org/10.7936/2yv8-zp85
Recommended Citation
Zhang, Heming, "Graph-AI Frameworks for Integrative Multi-Omics Data Analysis" (2026). Arts & Sciences Graduate Student Theses and Dissertations. 3841.
The definitive version is available at https://doi.org/10.7936/2yv8-zp85