KIT | KIT-Bibliothek | Impressum | Datenschutz

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

Sarmiento, Paulo Yanez; Rissom, Pia Francesca; Pfeuffer, Manuel; Simnacher, Marco; Safer, Jordan F.; Iqbal, Sumaiya; Heyne, Henrike O.; Klein, Nadja ORCID iD icon 1; Renard, Bernhard Y.
1 Scientific Computing Center (SCC), Karlsruher Institut für Technologie (KIT)

Abstract:

Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical properties of the protein are driving the prediction, we leverage an orthogonal projection technique that removes linear effects of known tabular features from embeddings and extend it to high-order and interaction effects. In this way, we remove the effects of interpretable biochemical features from PLM embeddings. In an ablation study, we show that this leads to a decrease in performance for a downstream classifier trained only on the embeddings to predict protein fitness. In an additional evaluation, we find that these biochemical features explain a substantial part of the variance in the predictions of this classifier. Hence, we can show that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness. ... mehr


Volltext §
DOI: 10.5445/IR/1000196703
Veröffentlicht am 31.08.2026
Cover der Publikation
Zugehörige Institution(en) am KIT Scientific Computing Center (SCC)
Publikationstyp Forschungsbericht/Preprint
Publikationsjahr 2026
Sprache Englisch
Identifikator KITopen-ID: 1000196703
HGF-Programm 46.21.02 (POF IV, LK 01) Cross-Domain ATMLs and Research Groups
Verlag arxiv
Umfang 17 S.
Externe Relationen Siehe auch
Schlagwörter explainable artificial intelligence (XAI) · interpretability · protein, language model (PLM) · embedding · orthogonal projection · protein fitness
Nachgewiesen in arXiv
KIT – Die Universität in der Helmholtz-Gemeinschaft
KITopen Landing Page