Designing Proteins Beyond Natural Sequences

Have you ever wondered how nature constructs its most complex biological machinery? It all comes down to proteins. A protein’s ultimate function is dictated by its three-dimensional structure, which is determined by its unique sequence of amino acids—the fundamental building blocks of life.

When scientists set out to engineer custom proteins—such as therapeutic molecules designed to bind to disease-causing agents inside our cells—they typically rely on a two-step computational design process. First, they define the target structure. Then, a machine learning algorithm generates a selection of amino acid sequences predicted to fold into that specific shape.

However, nature is incredibly flexible. In the natural world, many completely different amino acid sequences can fold into the exact same 3D structure. Conversely, a single sequence can adopt various shapes depending on environmental triggers or its inherent flexibility. Therefore, the grand challenge in using artificial intelligence for protein design is training neural networks to recognize this biological versatility—helping them “see” that there are multiple correct pathways to achieve the same structural fold.

“For years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select — our work shows that this isn’t the best metric for protein design,” explains Amy E. Keating, Department of Biology head, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of a paper recently published in PNAS.

To overcome this limitation, researchers have developed PottsMPNN, a groundbreaking machine learning framework. By integrating the physical laws that govern protein structure and molecular stability, PottsMPNN significantly improves sequence generation. It also boasts an enhanced ability to predict how mutations alter a protein’s structural integrity. Simply put, this new model possesses a much deeper understanding of the “sequence-energy landscape”—the delicate relationship between individual amino acids and the overall stability of the fold.

Integrating this framework into protein engineering pipelines will empower researchers to design entirely novel, structurally stable proteins with sequences that look nothing like anything found in nature.

“If we’re thinking about a completely novel, designed structure, there would be no native sequence to compare it to,” says graduate student and lead author Foster Birnbaum. “What we actually care about is how likely the generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it can predict the effect of mutations on the stability of the protein.”

Beyond the Noise: Revolutionizing Machine Learning in Biology

Just as generative AI has transformed creative industries, machine learning is rapidly accelerating the pace of fundamental biological discoveries. Reliable computational models for generating synthetic protein structures and sequences are a relatively recent breakthrough. Yet, one of the most widely used models in the scientific community today was released in 2022.

“For a field that’s moving as fast as machine learning in biology, that model has not been surpassed — we’ve been trying to understand why that is, and what it is about that model that makes it so useful,” Birnbaum notes.

To unlock new levels of performance, Birnbaum investigated the strategic application of “noise”—the process of introducing slight structural variations into the training data. This intentional variation prevents the AI model from simply memorizing and mimicking existing native sequences, thereby encouraging the generation of highly diverse structures.

Furthermore, PottsMPNN utilizes pairwise distribution to map the intricate physical interactions between amino acids. By calculating how all 20 possible amino acids interact at any given pair of positions within a protein, PottsMPNN maps the sequence-energy landscape with unprecedented accuracy compared to older methods.

Finally, the team introduced groups of evolutionarily related sequences into the training process of PottsMPNN. This step taught the framework a crucial lesson: how diverse evolutionary pathways can still converge on the same folded shape.

Birnbaum points out that while relying on evolutionary data might seem like a step back toward native sequences, PottsMPNN proves the opposite. As the model reduces its dependence on exact native sequences, its capacity to predict structural compatibility and energy stability for entirely novel proteins actually improves.

The Future of Custom Protein Design

“Once we can design any protein we want, that enables us to do a potentially scary amount of biological engineering,” says Birnbaum. “It’s a difficult task, but I’m really optimistic about this century’s progress in biology.”

Looking ahead, Birnbaum hopes the framework can be further refined and tailored to specific tasks, which has historically led to even more precise predictions regarding the outcomes of genetic mutations.

Ultimately, as Keating concludes, “Our methods move the field toward designing useful new-to-nature proteins for diverse applications while providing a stronger foundation for future advances.”

Photo credit & article inspired by: Massachusetts Institute of Technology

Leave a Reply

Your email address will not be published. Required fields are marked *