Award Date

5-15-2026

Degree Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

Department

Computer Science

First Committee Member

Mingon Kang

Second Committee Member

Fatma Nasoz

Third Committee Member

Bryar Shareef

Fourth Committee Member

Laxmi Gewali

Fifth Committee Member

Qian Liu

Number of Pages

96

Abstract

Trustworthy prediction of enzyme function from protein sequences remains a central challenge in computational biology, particularly when annotated data are limited, imbalanced, or incomplete. This dissertation develops interpretable deep learning methods for enzyme discovery and enzyme function prediction from amino acid sequences. First, it introduces PEPIC, an interpretable convolutional neural network for substrate-level prediction of hydrolytic plastic-degrading enzymes. Using curated and expanded sequence datasets, PEPIC improved predictive performance over benchmark methods, identified sequence regions aligned with catalytic and substrate-binding residues, and supported the discovery and experimental validation of a previously uncharacterized PET-degrading enzyme. Second, this dissertation investigates the integration of Kolmogorov-Arnold Networks (KANs) into state-of-the-art models for Enzyme Commission (EC) number prediction. KAN modules consistently improved predictive performance across multiple architectures, and a dedicated interpretation method recovered biologically meaningful motif sites from enzyme sequences. Third, it develops HIT-EC, a hierarchical interpretable transformer for EC number prediction that incorporates incompletely annotated sequences through a masked loss objective. Evaluated across repeated hold-out experiments, external benchmark data, and microbial genome analyses, HIT-EC consistently outperformed current state-of-the-art methods while providing trustworthy, biologically grounded evidence for its predictions. Taken together, these studies show that accurate enzyme annotation requires not only strong predictive performance but also interpretable and reliable evidence that can guide biological validation. This work advances sequence-based deep learning strategies for enzyme discovery and EC number prediction, and provides a practical foundation for more transparent computational annotation in enzymology and proteomics.

Controlled Subject

Computational biology; Enzymes--Analysis; Amino acid sequence

Disciplines

Biotechnology | Computer Sciences | Physical Sciences and Mathematics

File Format

PDF

File Size

23600 KB

Degree Grantor

University of Nevada, Las Vegas

Language

English

Rights

IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/


Share

COinS