dc.title: Two studies in molecular and computational biology: ATM loss driving transcriptional disruption, and the practical limits of zero-shot evolutionary models dc.description.abstract: This dissertation contains two independent bodies of work in molecular biology and computational biology. The first investigates the mechanistic basis of cerebellar neurodegeneration in Ataxia-Telangiectasia (A-T), a childhood-onset disorder caused by loss of the ATM kinase. The canonical framework (in which ATM's role in DNA double-strand break repair underlies all aspects of the disease) fails to explain why cerebellar Purkinje neurons are selectively vulnerable despite being post-mitotic and non-proliferative. Prior work has implicated a distinct, ROS-dependent ATM activation pathway that is genetically separable from its DSB repair function. This pathway is more specifically relevant to cerebellar integrity. Using RNA sequencing of A-T patient cerebellum and genome-wide mapping of poly(ADP-ribose) and RNA-DNA hybrids in human neuron-like cells, I provide additional evidence for this framework I show that ATM loss produces transcriptional dysregulation and transcription-coupled DNA damage accumulation that converge on highly transcribed GC-rich loci. These findings connect chromatin disruption and innate immune signaling to neurodegeneration and argue that transcriptional consequences of ATM loss, rather than DSB repair deficiency alone, represent a primary driver of cerebellar neurodegeneration in A-T.
The second body of work examines the practical limits of zero-shot protein mutation prediction. Zero-shot models trained on evolutionary sequence data are thought to predict the fitness of mutations without requiring protein-specific experimental training data. Through reanalysis of existing ProteinGym benchmarks and targeted controlled comparisons, I show that aggregate performance metrics systematically obscure substantial dataset-to-dataset variability. Furthermore, this variability tracks with the degree to which measured phenotypes correlate with the evolutionary signal the models encode. Model performance degrades predictably as phenotypes diverge from what natural selection has optimized. Namely, the discriminative power on gain-of-function mutations is negligible and models offer no reliable guidance for new-to-nature functional variants. These findings argue for considerably more skepticism toward existing zero-shot model evaluations, particularly in the clinical and protein engineering contexts where such models are most consequentially applied.

