Faculty Sponsor

Sharlee Climer

Final Abstract for URS Program

Large-scale genetic datasets from diverse diseases, including cancer and neurodegenerative disorders such as Alzheimer’s disease, provide unprecedented opportunities to study disease pathogenesis and generate hypotheses for drug discovery. However, identifying meaningful genetic interaction patterns within these datasets remains challenging due to combinatorial relationships amongst genes and heterogeneous biological underpinnings across subtypes. This project systematically analyzes a diverse collection of omics datasets (e.g., GSE44076, GSE97810, GSE127711, GSE112679) obtained from the NCBI Gene Expression Omnibus (GEO), spanning multiple biological conditions and data types. We use a unified computational pipeline designed to handle variation in dataset size, structure, and context. The steps in the pipeline include data preprocessing, normalization, case–control cohort splitting, gene interaction network construction using the DUO correlation metric, which captures heterogeneous gene relationships that traditional correlation-based approaches may overlook. Breadth-first search is then applied to identify connected components representing gene modules within the resulting networks. In colorectal cancer analysis, our initial run identified over 1000 significant patterns, which were then filtered across discovery and validation cohorts to retain 99 robust gene interaction patterns. In this presentation, we focus on results from colorectal cancer (GSE44076) and Alzheimer’s disease (GSE5281), where identified patterns demonstrate strong separation between disease and control samples. In colorectal cancer, many patterns show clear shifts in score distributions between tumor and normal samples, while the Alzheimer’s dataset reveals two distinct expression modules associated with disease progression. These findings demonstrate that heterogeneity-aware network modeling can effectively uncover coordinated gene expression patterns within large-scale genetic datasets and provide a powerful framework for discovering biologically meaningful genetic interactions and potential biomarkers for complex diseases.

Presentation Type

Oral Presentation

Document Type

Article

Share

COinS