Key result in one sentence. A collective variable learned from short wild-type simulations predicts which Chignolin mutations will accelerate or slow conformational transitions.
Abstract
While recent advances in AI have transformed protein structure prediction, protein function is also strongly influenced by the thermodynamic and kinetic features encoded in its underlying free-energy surface. Here, we propose a framework to rationally reshape this landscape in order to control conformational transition rates, built on the Collective Variables for Free Energy Surface Tailoring (CV-FEST) framework, and validate it on point mutations of the miniprotein Chignolin. The framework relies on Harmonic Linear Discriminant Analysis (HLDA)-based collective variables (CVs) constructed from short molecular dynamics trajectories confined to metastable basins, requiring only limited sampling within each basin. Notably, the HLDA CV derived solely from the wild-type system already provides residue-level scores that predict whether mutations at specific positions are likely to accelerate or slow unfolding transitions. Furthermore, we find that the leading HLDA eigenvalue associated with the derived CV, a quantitative measure of the one-dimensional statistical separation between folded and unfolded ensembles, is significantly correlated with transition rates across mutations. Together, these results suggest that kinetic effects of point mutations can be inferred from minimal local sampling, providing a practical route for guiding the engineering of transition rates without exhaustive simulations or large training data sets.
Overview
This work studies whether mutation-dependent peptide kinetics can be estimated without exhaustively simulating transition events for every candidate mutation. The approach uses molecular dynamics trajectories sampled within folded and unfolded metastable states to construct a collective variable that captures kinetic sensitivity.
The collective variable is constructed with Harmonic Linear Discriminant Analysis (HLDA) and used within the Collective Variables for Free Energy Surface Tailoring (CV-FEST) framework. The central assumption is that local fluctuations inside metastable states contain information about the free-energy barrier separating those states.
Results
For Chignolin point mutations, the HLDA collective variable derived from the wild-type system produces residue-level scores that identify positions where mutations are expected to accelerate or slow conformational transitions.
The leading HLDA eigenvalue, which measures one-dimensional statistical separation between folded and unfolded ensembles, also correlates with transition rates across mutations. This makes it a compact descriptor of mutation-dependent kinetic behavior.
Significance
Protein function is shaped not only by structure, but also by thermodynamics and kinetics: the relative stability of states, the height of free-energy barriers, and the rates of conformational transitions.
By estimating mutation effects from limited metastable-state sampling, the method supports a more efficient simulation-driven design loop: propose mutations, estimate their likely kinetic effect, and prioritize candidates before running more expensive rare-event simulations.
Technical Takeaways
The main technical point is not only that HLDA separates folded and unfolded states, but that this separation is predictive. The same low-dimensional description that distinguishes metastable ensembles also carries information about transition-rate changes under mutation.
That connects three things that are often treated separately:
- local equilibrium fluctuations inside metastable states
- free-energy barrier shaping through collective variables
- mutation-dependent conformational kinetics
Links
- Published paper
- PMC full text — PMCID: PMC13173532
- PubMed record — PMID: 42007551
- arXiv preprint
- DOI
- Code repository
- Dataset snapshot — version DOI: 10.5281/zenodo.18864706
- Dataset archive family — concept DOI: 10.5281/zenodo.18864705
Reproduction instructions
- Clone the code repository.
- Create the supplied Conda environment with
conda env create -f environment.yml, activateprotein-fes, then runpip install -e .. - Restore the archived data with
./scripts/unpack_data.sh. - Run the notebooks in
src/paper_plots/to regenerate the figures. The repository also documents how to regenerate the MFPT inputs from its PLUMED templates.
Cite the code
For the software and analysis workflow, cite this article and link to the reproduction repository. The repository’s CITATION.cff specifies the maintainers’ preferred citation.
Cite the dataset
For the exact archived snapshot used with this work, cite version DOI 10.5281/zenodo.18864706. Concept DOI 10.5281/zenodo.18864705 represents the archive across all versions and is the identifier cited in the published article’s data-availability statement.