EventsThe 4th International Electronic Conference on Catalysis Sciences
Published
This submission belongs to the session S4. Biocatalysis of the event The 4th International Electronic Conference on Catalysis Sciences
Published date
16 Sep, 2026
Academic Editor
author-avatarGonzalo De Gonzalo
Citation
Konstantinos Grigorakis, Zacharias Semertzidis, Evangelos Topakas, Pythia: A Surface-Based Deep Learning Framework for Predicting Carbohydrate-Binding Module Specificity, in Proceedings of The 4th International Electronic Conference on Catalysis Sciences, 22 September–24 September 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Pythia: A Surface-Based Deep Learning Framework for Predicting Carbohydrate-Binding Module Specificity

Zacharias Semertzidis 1
image
1. School of Chemical Engineering, National Technical University of Athens, Athens, Greece
Abstract

Carbohydrate-Binding Modules (CBMs) are non-catalytic domains found in many CAZymes, where they anchor the enzyme onto polysaccharide substrates and enhance activity. The specificity of a CBM, meaning which polymers it recognizes and binds, determines which substrates an enzyme bearing that CBM can effectively target, and is of growing interest for engineering biocatalysts active on materials ranging from cellulose to synthetic polymers. However, predicting this specificity from sequence alone remains unreliable, limiting the rational selection and engineering of CBMs for new substrate-targeting applications. Here, we present Pythia, a surface-based deep learning framework for predicting CBM–polymer binding specificity. Pythia represents each CBM as a two-dimensional surface property map using SURFMAP and applies convolutional neural networks (CNNs) to learn binding patterns directly from protein surface topology and physicochemical features, rather than sequence. Trained on a curated dataset of over 700 CBMs assayed against six polymers (Avicel, Chitin, Starch, PET, LDPE, PS), Pythia predicts CBM–polymer binding as a binary classification task, it achieves area-under-the-curve (AUC) values up to 0.71 on unseen CBMs, demonstrating that surface-derived features generalize substantially better than sequence-based descriptors in this data-limited setting. By making specificity prediction accessible directly from structural data, Pythia offers a scalable strategy for mining CBM databases for candidate substrate-targeting domains suitable for fusion with catalytic enzymes, supporting the design of biocatalysts with tailored substrate selectivity. Developed by the iGEM Athens 2025 team in collaboration with the IndBioCat group at the National Technical University of Athens, Pythia establishes a foundation for future CBM-based biocatalyst engineering, with direct relevance to applications such as microplastic-targeting enzyme systems, enzymatic plastic degradation and industrial biocatalysis.

Keywords
Carbohydrate-Binding Modules
Substrate specificity
Convolutional Neural Networks (CNNs)
Deep Learning
Protein engineering
Poster
ECCS2026_Grigorakis_poster.pdf
Green Catalytic Oxidation of Vanillyl Alcohol to Vanillin by Furoic Hydrazide-Derived Molybdenum Complexes
Tuning A/B- site cations in Gd-perovskite catalysts for selective CO2 hydrogenation to C2-C3 olefins: the role of Fe-Co synergy and potassium promotion