I. Literature Background
In 2023, the Nobel Laureate in Chemistry David Baker and his team published a review in Nature, systematically elaborating on the paradigm shift in de novo protein design—from physics-based computational methods to AI-driven generation—along with breakthrough advances in drug development, nanomachines, and related fields. This technology is expected to yield protein-based solutions for major diseases and environmental challenges within 5–10 years.
1.Molecular walkers
Molecular walkers lie at the fascinating crossroads of biology, chemistry, and physics, enabling the conversion of energy into unidirectional motion. Natural molecular motors, such as dynein, kinesin, and myosin, are essential for life processes, and although they have been studied in detail, our understanding of them remains incomplete. Programmable and robust de novo designed protein nanorobots could revolutionize fields ranging from precision medicine to materials. Despite recent remarkable progress in de novo protein design, building such functional systems remains elusive.
2.Current research status
Existing technologies include protein walkers assembled from natural DNA-binding proteins, which can move along DNA steps with the help of microfluidics and de novo designed rotor proteins. Tracking the motion of rotors has proven challenging. Designing a stochastic protein walker that can move over long distances would mark a substantial advance, deepen our understanding of molecular machine design, and lay the foundation for future protein-based nanodevices.
3.Research direction
Molecular machines have enormous potential. Great progress has been made in the design of static monomeric and oligomeric protein structures, but the design of dynamic protein systems has been limited.
Based on the above status, the Baker team published a research paper titled "De-novo design of a random protein walker" on the preprint platform bioRxiv In 2025, reporting the first de novo designed stochastic protein walking system. This breakthrough marks a key turning point in protein design, moving from static structural creation to dynamic functional construction, and provides a new technological platform for developing programmable "protein robots."

4.Advantages of this study
▶Binding specificity: The authors designed a completely de novo protein system consisting of 4‑ to 8‑foot walkers that specifically bind to designed protein tracks containing footholds.
▶Zero fuel consumption: Diffusion of these walkers along the track without fuel consumption was made possible and characterized, thereby exploring how multivalency affects speed and processivity. This minimal system provides a versatile scaffold into which external energy sources can be incorporated in future designs.
II. Research Results
1. Track design strategy and structural features
The authors reasoned that at least three components are required to build a stochastic walker: a long protein filament as the track, a symmetric protein scaffold as the walker, and well-behaved reversible heterodimers as the feet and footholds that form the key interface enabling the walker to move along the track.

Figure 1. Basic components of the stochastic walker system
Protein filaments assembled from globular monomers have been successfully designed, but they are too short to be clearly resolved by total internal reflection fluorescence (TIRF) microscopy, a key technique for single‑molecule tracking experiments. To address this limitation, the authors aimed to design filaments with larger diameters to increase their rigidity, further improving filament length and structural stability.
The core of the track was designed by docking protein scaffolds and redesigning interfaces. The authors developed Fibre H‑fuse, a new method for rigidly attaching arbitrary proteins to the filament assembly. Two of these proteins – HA4 and HA13 – were both expressed in soluble form and formed filaments, confirmed by confocal microscopy. The structure of HA13 was further characterized by cryoEM at 3.3 Å resolution, revealing a filament with an inner diameter of 10 nm and a longitudinal spacing of 4.7 nm between the connected DHRs.
In addition to facilitating electron microscopy, the fusion of DHR provided additional spacing from the filament interface, reducing the chance of perturbing filament assembly when adding footholds for the walker.

Figure 2‑1. Structural features of the track
2. Selection of heterodimeric foot / foothold pairs
To accomplish tracking, heterodimeric protomers that can bind and dissociate are required as footholds on the filament. Initial tests with the shorter heterodimeric four‑helix bundle mALb811 indicated that the presence of footholds interfered with filament assembly.

Figure 2‑2. Structural features of the track
3. Walker design
The authors attached the P4SN coiled‑coil foot to the scaffold via a flexible six‑residue linker; P4SN specifically binds to the complementary P3SN coiled‑coil foothold displayed on the filament.
All six walker constructs were successfully expressed in E. coli and purified from the soluble fraction. Oligomeric assembly states were validated using two size‑based methods: size‑exclusion chromatography coupled with multi‑angle light scattering (SEC‑MALS) and mass photometry (MP), performed at micromolar and nanomolar concentrations, respectively. Both techniques confirmed the expected molecular weights and oligomeric states of the designed assemblies.
Figure 3. Biophysical characterization and structural validation of designed walker oligomers
4. Single‑molecule tracking of walkers on the track
To investigate the dynamics of walkers along the track, the authors used TIRF microscopy and fluorescence microscopy. Label‑free imaging was achieved by visualizing tracking in the green channel using TIRF. Walkers were labeled with ATTO 643 and imaged in the far‑red channel by TIRF at a frame rate of 4 Hz. Walker positions were determined by fitting the center of the fluorescence intensity distribution using Picasso 17. From the walker trajectories, the distribution of step lengths was extracted and fitted to a two‑dimensional Gaussian function.
Diffusion was measured for walkers with different numbers of feet along the remaining footprints. Two main trends emerged: walkers moved faster on tracks with higher‑affinity footholds (P3SN_4h) compared to those with lower binding affinity (P3SN_3.5h), and diffusion rates increased with walker valency.
To assess whether mNeonGreen attached to the track affected walker mobility, diffusion of WALKER‑C4 and WALKER‑C8 was measured on tracks lacking the fluorescent protein. Diffusion rates on tracks with and without mNeonGreen were comparable, indicating that the presence of the flexibly tethered fusion did not measurably affect walker movement once bound to the track.
The authors also evaluated the effects of temperature, viscosity, and salt concentration on the movement of WALKER‑C8 on tracks lacking mNeonGreen. At 25 °C (1.695 cP) the viscosity did not slow down the walkers; in 50 % sucrose, the diffusion rate decreased approximately 2‑fold on both tracks. At 500 mM salt, the diffusion rate also decreased approximately 1.5‑fold, most likely due to reduced affinity between the foot and the foothold. Conversely, when the temperature was increased from 25 °C to 45 °C, their diffusion rate increased approximately 1.5‑fold.

Figure 4. Single‑molecule tracking and diffusion analysis of engineered stochastic protein walkers
Research on recombinant protein design continues to be a hot topic. KMD Bioscience offers a variety of customized services including protein expression, recombinant protein design, recombinant protein preparation, and protein purification to meet customer needs at different experimental stages.
Q1:What problems during the design phase can lead to protein expression failure?
A1:
Codon bias mismatch: If the exogenous gene contains codons that are rare in the host, translation can be severely slowed. To address this, the gene sequence can be codon‑optimized, or specialized strains such as Rosetta that supply rare tRNAs can be used.
Signal peptide issues: Using the full‑length sequence of a eukaryotic protein (including its signal peptide) directly for prokaryotic expression can lead to misfolding. The signal peptide should be removed before cloning.
Improper tag selection: An N‑terminal tag may be cleaved off due to the signal peptide. For proteins prone to degradation, the fusion tag can be placed at the N‑terminus; for secreted proteins, a C‑terminal tag is recommended.
Q2:What problems during the expression phase can lead to protein expression failure?
A2:
Low expression yield: This is the most common difficulty. Besides optimizing the gene and expression system, induction conditions (timing, temperature, inducer concentration) should be adjusted, and different lysis buffer concentrations can be tried.
Inclusion body formation: When protein is overexpressed or expressed too rapidly, it may aggregate into inactive "inclusion bodies" without proper folding. This can be improved by low‑temperature slow induction (e.g., 25 °C), using a weak promoter, or adding osmotic regulators such as sorbitol to the medium. For prokaryotic hosts, switching to strains that promote disulfide bond formation, such as Origami, may also help.
Protein degradation: Proteases act as cellular "scavengers" and can attack exogenous proteins. The core solutions are to “reduce exposure” and “inhibit activity”: shorten the culture time after induction, add protease inhibitors to the lysis buffer, or switch to host strains lacking protease genes (e.g., lon, ompT).
Insufficient post‑translational modifications: Protein function often depends on “late‑stage modifications” such as glycosylation or phosphorylation. Insufficiency often indicates insufficient host capability. To solve this, one can upgrade from a simple E. coli system to yeast, insect, or even mammalian cell systems.
Q3:What is the workflow for recombinant protein production?
A3:Gene cloning and vector construction: Insert the gene of interest into an expression vector equipped with elements such as a promoter and a tag (e.g., His tag).
Host transformation and expression screening: Introduce the vector into the host cells and screen for positive clones through small‑scale culture.
Scale‑up culture and induction of expression: After successful screening, scale up the culture and optimize parameters such as induction time, temperature, and inducer concentration.
Cell lysis and sample clarification: Harvest the cells, lyse them by methods such as sonication or high‑pressure homogenization, and centrifuge to obtain the supernatant containing the target protein.
Protein purification: Use multi‑step purification strategies such as affinity chromatography and ion‑exchange chromatography to obtain high‑purity protein.
Q4:What are the selection principles and advantages of different expression systems?
A4:
Expression System | Advantages | Disadvantages | Best Application Scenarios |
E. coli | Low cost, fast turnaround, high yield, simple operation | Prone to inclusion bodies, lacks posttranslational modifications | Rapidly obtaining large amounts of protein when no modifications are required |
Yeast | Combines prokaryotic operability with eukaryotic folding/modification, secretable expression | Hyperglycosylation issues, glycan patterns differ from human | When some glycosylation is needed and cost is sensitive; Pichia pastoris is most widely used |
Insect cells | Large gene capacity, high soluble protein fraction, suitable for toxic proteins | Long turnaround, high cost, incomplete glycosylation | Expressing complex proteins, toxic proteins, or requiring more advanced eukaryotic processing |
Mammalian cells | Highest activity, achieves humanlike glycosylation closest to native | Long turnaround, very high cost, relatively low yield | Applications requiring fully humanized proteins, such as drug discovery and antibody production – considered the “gold standard” |
Q5:What problems may arise after protein purification?
A5:
Low purity after purification: Often caused by ineffective exposure of the His tag or insufficient washing. This can be improved by adding low concentrations of imidazole, glycerol, or adjusting the pH in the buffer.
Poor protein solubility: Possible reasons include the solution pH being close to the protein’s isoelectric point (leading to precipitation) or aggregation due to high concentration. Optimize buffer composition (pH, ionic strength), lower the protein concentration, and add carrier proteins such as BSA to low‑concentration working solutions to improve stability.
Low or no protein activity: Often caused by failed refolding or improper storage. Optimize inclusion body refolding conditions, or immediately aliquot and freeze to avoid repeated freeze‑thaw cycles.
0