Machine learning-driven simulations of the SARS-CoV-2 fitness landscape from deep mutational scanning experiments

Predicting protein - variant effects is a key challenge in preparing - for pathogenic viral strains, understanding mutation-linked diseases, and designing new proteins. Protein sequence-structure-function - relationships are difficult to model due to complex allosteric and epistatic effects. To inve...

Description complète

Enregistré dans:
Détails bibliographiques
Auteurs principaux: Durumeric, Aleksander (Auteur) , McCarty, Sean (Auteur) , Smith, Jay (Auteur) , Köhler, Jonas (Auteur) , Elez, Katarina (Auteur) , Raich, Lluís (Auteur) , Suriana, Patricia A. (Auteur) , Sztain, Terra (Auteur)
Format: Article (Journal)
Langue:anglais
Publié: May 25, 2026
In: Journal of chemical information and modeling
Year: 2026, Volume: 66, Numéro: 10, Pages: 5721-5735
ISSN:1549-960X
DOI:10.1021/acs.jcim.6c00332
Accès en ligne:Verlag, kostenfrei, Volltext: https://doi.org/10.1021/acs.jcim.6c00332
Accéder au texte intégral
Notes sur l'auteur:Aleksander E.P. Durumeric, Sean McCarty, Jay Smith, Jonas Köhler, Katarina Elez, Lluís Raich, Patricia A. Suriana, and Terra Sztain
Description
Résumé:Predicting protein - variant effects is a key challenge in preparing - for pathogenic viral strains, understanding mutation-linked diseases, and designing new proteins. Protein sequence-structure-function - relationships are difficult to model due to complex allosteric and epistatic effects. To investigate efficient modeling strategies, we trained supervised machine learning (ML) models with deep mutational scanning (DMS) libraries of SARS-CoV-2 receptor binding domain (RBD) sequences labeled with angiotensin converting enzyme 2 (ACE2) binding affinity. These models demonstrate superior performance predicting combinatorial mutation effects compared to adding or averaging the effects of point mutations and exhibit strong extrapolative performance ranking omicron variants when training only near wild type (WT) variants. We characterize the RBD fitness landscape by combining ML with Markov Chain Monte Carlo simulations to predict evolutionary patterns from the WT sequence. These generate comparable sequence profiles to high-fitness sequences in DMS data and predict mutations in unseen omicron variants. These models provide insight into the relationship between RBD sequence elements and offer a new perspective on the use of DMS to predict emerging viral strains, which we anticipate will be applicable to other evolutionary prediction tasks. To facilitate application and - future development of this strategy, we introduce Mavenets: https://github.com/SztainLab/mavenets.
Description:Online veröffentlicht: 06. Mai 2026
Gesehen am 10.08.2026
Description matérielle:Online Resource
ISSN:1549-960X
DOI:10.1021/acs.jcim.6c00332