Skip to content
Community

Protein Backbone to Sequence with RFdiffusion and ProteinMPNN

Generate a small monomer backbone with RFdiffusion, pass its PDB coordinates to ProteinMPNN, and retain paired structures, sequences and provenance for later validation.

Level: Advanced Cost: Free Privacy: Local ~90 min
Start Setup

You'll be able to

  • Produce traceable computational candidate artifacts and inspect the documented acceptance criteria.

What you'll build

Documentation-based review draft. This workflow has not been executed in this batch; no installation, inference, optimization or experimental result has been verified.

Scope and audience

For computational protein-design researchers. The baseline is an unconditional 150-residue monomer, not a binder or functional protein claim. RFdiffusion generates backbone coordinates; glycine labels in designed regions are placeholders. ProteinMPNN assigns sequences to an inspected backbone. No FastRelax, refolding or experimental validation is included.

Inputs, handoff and artifacts

Input: a length constraint, selected RFdiffusion checkpoint and exact source/environment records. Check the resulting PDB contains one chain A and usable N/CA/C/O coordinates for the ordinary full-backbone ProteinMPNN model. Preserve chain IDs, residue numbering and the RFdiffusion PDB/TRB pairing; gaps, missing atoms, alternate locations and constrained motifs require an explicit mapping rather than blind conversion. The TRB file may contain serialized Python objects: inspect it only from a trusted run. Output: backbone PDB and metadata, FASTA in the ProteinMPNN seqs directory, sampling seed/temperature and score records. Scores rank model preferences, not activity or folding success.

Limits, costs and terms

RFdiffusion's inspected environment is Python 3.9, PyTorch 1.9 and CUDA 11.1; modern GPU compatibility is not established here. Keep ProteinMPNN in a separately recorded environment and exchange files. GPU time, storage and dependency setup are user costs; the setup estimate excludes inference. Code licenses are BSD-3-Clause and MIT respectively; retain notices and check the selected checkpoint's accompanying terms before use. Do not infer rights to private input structures from software licensing. No broad OS or hardware guarantee is made.

Official references

Backbone generation

RFdiffusion

Backbone-conditioned sequence design

ProteinMPNN

Stack Components

RFdiffusion

Backbone generation · source 86507b6538f51fce57b5a72477165f03999ed7ae; Base_ckpt.pt

Documentation-based review draft. This workflow has not been executed in this batch; no installation, inference, optimization or experimental result has been verified.

Code is open source; compute/storage and any external-service conditions remain the user's responsibility.

View Resource

ProteinMPNN

Backbone-conditioned sequence design · source 8907e6671bfbfc92303b5f79c4b5e6ce47cdef57; v_48_020

Documentation-based review draft. This workflow has not been executed in this batch; no installation, inference, optimization or experimental result has been verified.

Code is open source; compute/storage and any external-service conditions remain the user's responsibility.

View Resource

Compatibility

ClientOSArchitectureVersion requirements
Python LinuxAny>= 3.9

Setup & Test

1. Define a bounded monomer design

Linux

Choose an unconditional 150-residue monomer for the first check. Obtain authorized inputs, storage, compatible compute and the official checkpoint; the original CUDA environment needs its own installation review.

Official source

Expected result

A written objective and recorded checkpoint/environment prerequisites.

2. Obtain the inspected source versions

Linux

Run from a new working directory. Follow the pinned RFdiffusion README's SE3nv/SE3-Transformer installation and checkpoint placement; separately install ProteinMPNN's documented Python/PyTorch/NumPy requirements. These clone commands alone are not a complete installation.

git clone https://github.com/RosettaCommons/RFdiffusion.git
git -C RFdiffusion checkout 86507b6538f51fce57b5a72477165f03999ed7ae
git clone https://github.com/dauparas/ProteinMPNN.git
git -C ProteinMPNN checkout 8907e6671bfbfc92303b5f79c4b5e6ce47cdef57
Official source

Expected result

Pinned repositories and separately recorded, dependency-complete environments.

3. Record the file and chain contract

Linux

Use the full-backbone v_48_020 model, not --ca_only. Keep chain A and inspect N/CA/C/O atoms and residue correspondence. If a motif or additional chain is introduced, add an explicit fixed-position/chain policy before redesigning.

Official source

Expected result

An inspected single-chain PDB can be passed to the documented --pdb_path interface.

4. Generate one candidate backbone

Linux

Activate the configured RFdiffusion environment first. Run a single small sample; inspect the resulting files rather than treating command completion as scientific success.

cd RFdiffusion
./scripts/run_inference.py 'contigmap.contigs=[150-150]' inference.output_prefix=review_outputs/backbone inference.num_designs=1
Official source

Expected result

Expected artifact: RFdiffusion/review_outputs/backbone_0.pdb and corresponding TRB metadata, with 150 designed residues; this is not an observed result.

5. Design sequences from the inspected PDB

Linux

Return to the parent working directory and activate the ProteinMPNN environment. Pass the actual output PDB only after chain/atom checks; use seed 37 and temperature 0.1 as explicitly chosen demonstration settings.

python ProteinMPNN/protein_mpnn_run.py --pdb_path RFdiffusion/review_outputs/backbone_0.pdb --pdb_path_chains A --out_folder sequence_outputs --num_seq_per_target 2 --sampling_temp 0.1 --seed 37 --model_name v_48_020
Official source

Expected result

Expected: FASTA under sequence_outputs/seqs with sampled sequences and model/seed/score provenance. Verify sampled sequence count and 150-residue length; the reference entry is not a new design.

Troubleshooting

  • Dependency or CUDA failure: use the pinned upstream environment and review the actual driver/device; do not silently mix environments.
  • Missing backbone: inspect RFdiffusion logs and checkpoint placement.
  • Parser errors: check chain A, residue numbering and complete backbone atoms; preserve mappings when repairing inputs.
  • Memory limits: reduce design batch size and retain the error.
  • FASTA mismatch: distinguish reference from sampled sequences; match every sequence to its backbone. Low model score alone is not validation.
Still not working

Alternatives

A separately reviewed ProteinMPNN-FastRelax protocol is an extension, not an equivalent executed step. This baseline deliberately stops at backbone/sequence pairs.