Date of Award

8-2026

Document Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

Department

Computer Science

Committee Chair/Advisor

Feng Luo

Committee Member

Rong Ge

Committee Member

Kai Liu

Committee Member

Yongkai Wu

Abstract

The advancement of single cell RNA sequencing (scRNA-seq) has enabled the study of causal relationships between genes at single cell resolution. Although many causal discovery methods have been applied to scRNA-seq perturbation data, they are not well-suited to capture the characteristics of scRNA-seq data. The overall goal of this dissertation is to enhance researchers' ability to gain insight into genetic relationships.

The scRNA-seq data is high-dimensional, typically containing thousands of genes, and is sparse and zero-inflated due to dropout events, as well as noisy and subject to biological variability. Traditional causal discovery methods, such as constraint or score-based approaches, do not scale well to high-dimensional settings. In our first study, we address these drawbacks with a zero-inflated negative binomial (ZINB) approach that explicitly models the data-generating process and is better suited to the characteristics of single cell data. In the first study, we combine causal inference methods with ZINB to simulated and real-world single cell datasets with perturbations. We demonstrate that the ZINB model outperforms Gaussian alternatives in reconstructing gene expression profiles. Biological enrichment analysis of the resulting causal graphs supports that our method captures meaningful regulatory pathways and known genes linked to immune system and tumor interactions.

In the second study, we present a novel causal clustering framework for scRNA-seq data. We introduce a causal merge step that evaluates similarities between clusters based on their respective causal gene regulatory networks. This enables the identification of biologically meaningful cell clusters. We compare our approach against state-of-the-art single cell clustering methods on real-world single cell datasets, demonstrating competitive clustering performance and improved detection of rare cellular populations. Finally, we perform causal discovery within rare-cell groups and conduct biological enrichment analysis on the resulting causal graphs to further assess the biological relevance of the discovered clusters.

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.