Original Authors: Yang W, Wang S, Lee GR, et al. (David Baker Group)
Published in: Nature, 2026
Abstract
Protein design is undergoing a paradigm shift from imitating nature to surpassing it. Based on a recent review by Yang et al. from David Baker's group published in Nature, this article systematically reviews the technical evolution of de novo protein design, analyzes the capability boundaries of current methods, and explores the challenges and future directions in this field.
1. Introduction: The Paradigm Shift in Protein Design
Proteins are the core executors of life activities, with their functions determined by three-dimensional structures. Traditional protein engineering relies on modifying natural proteins, while de novo protein design attempts to build entirely new proteins with specific functions from scratch. The development of this field has undergone an important transformation from physics-driven to data-driven approaches.
2. Historical Evolution: Three Development Stages
2.1 The Physics-Based Era
The 1990s to 2010s marked the physics-based era of protein design. Early protein design was primarily based on energy minimization principles and physicochemical knowledge. Researchers predicted the most stable structures by calculating the energy of amino acid sequences in specific conformations. Key advances during this phase included:
- Development of the Rosetta platform, laying the foundation for protein structure prediction and design
- Successful design of simple topological structures such as α-helix bundles
- Deep understanding of protein folding thermodynamics principles
However, this period was characterized by high computational costs, low design success rates, and difficulty handling complex structures.
2.2 The Data-Driven Era
From the 2010s to 2020s, with the accumulation of protein structure databases and the development of deep learning technology, protein design entered the data-driven stage:
- Deep learning models such as AlphaFold achieved breakthroughs in structure prediction
- Design methods based on large-scale sequence and structure data emerged
- Important progress was made in protein-protein interaction design
However, model interpretability was insufficient, and the ability to handle out-of-distribution cases was limited.
2.3 The Generative AI Era
From the 2020s to the present, generative AI represented by diffusion models and large language models is reshaping protein design:
- Methods such as RFdiffusion have achieved unconditional protein generation
- Function-oriented design has become possible, including enzyme, vaccine, and drug delivery carrier design
- Experimental validation success rates have significantly improved
3. Core Technical Methods
3.1 Structure Prediction-Driven Design
Deep learning-based structure prediction models such as AlphaFold and ESMFold provide important tools for design:
- Sequence Design: Given a target structure, optimize the amino acid sequence to stabilize it
- Structure Generation: Generate reasonable protein backbones from noise
- Joint Sequence-Structure Optimization: Simultaneously optimize sequences and structures to achieve design goals
3.2 Diffusion Models in Protein Design
Diffusion models achieve high-quality protein structure generation by learning the inverse process from noise to data. This method can generate diverse protein topologies, support conditional generation such as specifying functional sites, and show potential in drug delivery carrier and enzyme design.
3.3 Functional Design Strategies
The leap from structure to function is the current core challenge:
- Active Site Design: Embedding catalytic or binding sites into protein backbones
- Protein-Protein Interaction Design: Designing proteins with specific binding characteristics
- Dynamic Property Design: Controlling protein conformational changes and allosteric effects
4. Application Cases and Progress
4.1 Vaccine Design
De novo designed protein nanoparticles can serve as vaccine carriers, displaying antigen arrays to elicit strong immune responses. During the COVID-19 pandemic, such designs entered clinical trials, validating the feasibility of this strategy.
4.2 Enzyme Design
Designing proteins with catalytic activity is the holy grail of the field. Current progress includes:
- Catalyst design for simple reactions such as ester hydrolysis
- Design of metal ion-dependent enzymes
- Reaction specificity design based on transition state theory
However, enzyme design for complex multi-step reactions remains challenging.
4.3 Drug Delivery Carriers
Designing self-assembling protein nanocontainers for targeted delivery of small molecule drugs or nucleic acids requires:
- Controllable assembly and disassembly properties
- Precise arrangement of targeting ligands
- Biocompatibility and immunogenicity optimization
4.4 Protein Switches and Sensors
Designing protein switches that respond to specific signals such as small molecules, temperature, or pH, including:
- Allosteric switch design
- Fluorescent protein sensors
- Logic-gated protein systems
5. Current Challenges and Limitations
5.1 The Design-Function Gap
The ability to design stably folded protein structures does not mean achieving the intended biological functions. The mapping from structure to function remains an incompletely solved problem.
5.2 Dynamic Property Design
Protein functions often depend on their dynamic properties such as conformational changes and flexible regions. Current methods have limited control over dynamic properties.
5.3 Experimental Validation Bottleneck
Computationally designed proteins still require experimental validation. Advances in high-throughput screening and characterization technologies are key to accelerating the design-validation cycle.
5.4 Membrane Protein Design
Membrane proteins occupy an important position in drug targets, but their design difficulty is much higher than soluble proteins, mainly limited by:
- Complexity of membrane environment simulation
- Technical difficulty of experimental characterization
- Scarcity of structural data
6. Future Outlook
6.1 Deep Integration with Experimental Technologies
Next-generation protein design will be more tightly integrated with high-throughput synthesis and screening technologies, structural characterization methods such as cryo-EM, and machine learning-assisted experimental design.
6.2 Multi-Scale Modeling
Cross-scale modeling from atomic-level interactions to cellular-level functions will support more complex protein system design.
6.3 Further Integration of Artificial Intelligence
Technologies such as large language models, multimodal learning, and reinforcement learning are expected to further enhance protein design capabilities.
6.4 Application Expansion
Protein design will play a greater role in the following areas:
- Sustainable Chemistry: Biocatalysis
- Precision Medicine: Targeted therapy
- Biomaterials: Self-assembling materials
- Environmental Remediation: Pollutant degradation
7. Conclusion
De novo protein design is moving from academic exploration to practical application. Although the leap from structure to function remains challenging, the introduction of generative AI has significantly accelerated the development of this field. For practitioners, understanding the applicable boundaries of methods, establishing reasonable design expectations, and closely integrating with experimental validation are key to achieving successful designs.
References
Yang W, Wang S, Lee GR, et al. The past, present and future of de novo protein design. Nature. 2026; doi: 10.1038/s41586-026-10328-7
← Back to BlogThis article is based on objective analysis of academic literature and does not constitute any investment or R&D advice.