Award Date
5-15-2026
Degree Type
Dissertation
Degree Name
Doctor of Philosophy (PhD)
Department
Computer Science
First Committee Member
Mingon Kang
Second Committee Member
Fatma Nasoz
Third Committee Member
Bryar Shareef
Fourth Committee Member
Laxmi Gewali
Fifth Committee Member
Qian Liu
Number of Pages
96
Abstract
Trustworthy prediction of enzyme function from protein sequences remains a central challenge in computational biology, particularly when annotated data are limited, imbalanced, or incomplete. This dissertation develops interpretable deep learning methods for enzyme discovery and enzyme function prediction from amino acid sequences. First, it introduces PEPIC, an interpretable convolutional neural network for substrate-level prediction of hydrolytic plastic-degrading enzymes. Using curated and expanded sequence datasets, PEPIC improved predictive performance over benchmark methods, identified sequence regions aligned with catalytic and substrate-binding residues, and supported the discovery and experimental validation of a previously uncharacterized PET-degrading enzyme. Second, this dissertation investigates the integration of Kolmogorov-Arnold Networks (KANs) into state-of-the-art models for Enzyme Commission (EC) number prediction. KAN modules consistently improved predictive performance across multiple architectures, and a dedicated interpretation method recovered biologically meaningful motif sites from enzyme sequences. Third, it develops HIT-EC, a hierarchical interpretable transformer for EC number prediction that incorporates incompletely annotated sequences through a masked loss objective. Evaluated across repeated hold-out experiments, external benchmark data, and microbial genome analyses, HIT-EC consistently outperformed current state-of-the-art methods while providing trustworthy, biologically grounded evidence for its predictions. Taken together, these studies show that accurate enzyme annotation requires not only strong predictive performance but also interpretable and reliable evidence that can guide biological validation. This work advances sequence-based deep learning strategies for enzyme discovery and EC number prediction, and provides a practical foundation for more transparent computational annotation in enzymology and proteomics.
Controlled Subject
Computational biology; Enzymes--Analysis; Amino acid sequence
Disciplines
Biotechnology | Computer Sciences | Physical Sciences and Mathematics
File Format
File Size
23600 KB
Degree Grantor
University of Nevada, Las Vegas
Language
English
Repository Citation
Dumontet, Louis, "Interpretable Deep Learning Models for Trustworthy Prediction of Enzyme Functions" (2026). UNLV Theses, Dissertations, Professional Papers, and Capstones. 5534.
https://oasis.library.unlv.edu/thesesdissertations/5534
Rights
IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/