Data from: A hybrid physics-deep learning framework for combinatorial de novo design of small-molecule binding proteins
Data files
May 06, 2026 version files 5.99 GB
-
README.md
11.04 KB
-
supplementary_data.zip
5.99 GB
Abstract
Engineering small-molecule binding proteins de novo remains a significant challenge as even advanced generative models struggle to model the atom-level details of protein-ligand interactions with sufficient accuracy. Higher experimental success rates have resulted from methods that explicitly scaffold predefined binding interactions into helical bundles. Here, we introduce a scaffolding strategy that generalizes to alpha-beta architectures. By screening thousands of combinatorially assembled protein-ligand interactions against diverse de novo backbones with finely varied pocket geometries, the protocol allows for high-fidelity accommodation of target interaction geometries. Our protocol then integrates physics-based and deep learning methods for optimization of interfacial interactions and sequence-structure compatibility, considerably improving in silico design metrics. Applying this method to two chemically similar steroids achieved a notable experimental success rate (4/26 designs bind their targets), and NMR structures of two designs are in good agreement with design models. Our generalizable, atomically precise approach offers a robust framework for small-molecule binder design, effectively eliminating the need for high-throughput screening.
Description of the data and file structure
supplementary_data.zip: Contains all relevant supplementary data, with folders therein described below.
Benzene_Clustered_Interactions: Data pertaining to Figure 1D. Contains a json file organized as a python dictionary where each key is an amino acid 3 letter code (eg LYS) and the corresponding value is a list of tuples (population, energy)representing, for each clustered interaction mode between the given amino acid and the chemical substructure (benzene), the population of the cluster and the two body rosetta energy score (fa_atr=1,fa_rep=.55,hbond_sc=1,fa_elec=1,hbond_bb_sc=1) of the amino acid-benzene interaction. See Methods: Motif library generation.
MOAD_Benchmark: Data pertaining to Supplementary Figure S4. Contains a JSON file organized as a Python list of dictionaries. Each dictionary corresponds with a different structure in the dataset, with each key specifying a score type and each value the actual score for that term. Additionally, the PDB code for each structure is provided under the key "PDBID", and the Kd for binding (nM) is provided under the key "KD_nm". Dataset curation and scoring parameters are described in Methods: Derivation of in silico filtering criteria.
Subsets_for_in_silico_success_rate_comparisons: Contains both the structures (.pdb files) and scores (.json files) for all CLAIRE-derived design subsets and RFDaa-derived designs used for comparisons of in silico success rates (Figure 1G, Figure 2A, 2B, 2C, 2D, Table S7, Table S8).
Each subdirectory contains data for a different design subset, specified in the subdirectory name according to [Target molecule 3 letter code]_[design method]. ie:
esl_unrefined_CLAIRE = CLAIRE derived estriol designs before refinement steps
esl_hbrefined_CLAIRE = CLAIRE derived estriol designs after application of HBRefine algorithm
esl_stability_refined_CLAIRE = CLAIRE derived estriol designs after stability-oritented refinement with ProteinMPNN
prg_unrefined_CLAIRE = CLAIRE derived progesterone designs before refinement steps
prg_hbrefined_CLAIRE = CLAIRE derived progesterone designs after application of HBRefine algorithm
prg_stability_refined_CLAIRE = CLAIRE derived progesterone designs after stability-oritented refinement with ProteinMPNN
esl_rfdaa = RosettaFold Diffusion all-atom derived designs for estriol
esl_rfdaa_hbrefined = RosettaFold Diffusion all-atom derived designs for estriol after application of HBRefine algorithm
prg_rfdaa = RosettaFold Diffusion all-atom derived designs for progesterone
prg_rfdaa_hbrefined = RosettaFold Diffusion all-atom derived designs for progesterone after application of HBRefine algorithm
The .json files are organized as a list of Python dictionaries. Each dictionary corresponds with a different structure in the dataset, with each key specifying a score type and each value the actual score for that term. The different scoring terms are described in Supplementary Table S1. The generation of these datasets is described in Methods: Design test sets for comparison of in silico success rates.
For several design subsets, as a consequence of distributed computing, the main subset directory contains either several subdirectories (e.g., esl_stability_refined_CLAIRE, subdirectories = batchX) across which design models and score JSON files are distributed, or else only the scores are distributed across several JSON files (eg esl_rfdaa, prg_rfdaa, prg_rfdaa_hbrefined, esl_rfdaa_hbrefined, JSON files = XXXX.json where XXXX is a distinct number).
For the subsets prg_rfdaa and esl_rfdaa, given the large number of designs, the design model .pdb files are included as a distinct, compressed directory, being 'prg_rfdaa_design_models.zip' and 'esl_rfdaa_design_models.zip' for progesterone and estriol designs, respectively.
LUCS_NTF2_scaffold_library: Contains .pdb files for all structures in our LUCS-generated NTF2 scaffold library (1,816 structures). The production of these structures is described in Methods: Protein scaffold library generation.
Ordered_DNA_Seqs: Contains DNA sequences for 13 designed estriol binders, 13 designed progesterone binders, and 5 design point mutants that were ordered for experimental testing as described in methods: Production runs, the estriol binders were ordered in two batches containing 9 and 4 designs, respectively. The sequences of the former 9 are described in a .csv file esl_CLAIRE_designs_platemap, while the latter 4 are listed in individual .txt files named according to the format esl_[design ID].txt.
PhiPsi_comparison: contains .json files with phipsi values for CLAIRE and RFDaa generated designs used for comparison of secondary structure content (Supplementary Figure S5). The data is organized as Python dictionaries where each key represents a different design ID, and each value is a list of two lists, with the former corresponding the phi values and the latter the psi values for each amino acid in the design. The ordering of the phi and psi list values are the same, such that the element in each at index i corresponds to the same amino acid.
boltz_2_predictions: contains boltz-2 outputs for experimentally tested CLAIRE designs, as seen in Supplementary Figure S7, Supplementary Table S4, and Supplementary Table S5. There is a directory for each target molecule (estriol/progesterone) containing subdirectories for each design, with these containing the boltz-2 predicted structure (.pdb), affinity and confidence metrics (.json), and error/confidence matrices (.npz)
af2_models: contains alphafold2 predictions for experimentally tested CLAIRE designs. There is a subdirectory for each target molecule (estriol/progesterone) containing .pdb files for the best predicted structure for each design, along with a JSON file organized as a Python dictionary of dictionaries, where each key is a design ID, and the corresponding value is a dictionary with the average and per-residue lddt values for each of the 5 af2 models.
CLAIRE_design_models: contains final, refined design models (.pdb files) for experimentally characterized CLAIRE designs targeting both estriol and progesterone.
design_with_naturally_occurring_proteins: contains inputs, matches, and designs obtained from applying the design workflow to naturally occurring motifs and NTF2 scaffolds present in the protein data bank. See Methods: Matching with motifs and scaffolds available in PDB. There are 3 subdirectories:
natural_ntf2_library: contains the scaffold structures (.pdb) and position files (.pos) used for matching
esl_nat_matching: contains inputs: params files (.params) and motifs (.pdb,.cst) used for matching, as well as the source structures (.pdb) from which the motifs were derived.
also contains matches obtained (.pdb) and subsequent designs, .resfiles used for design, and rosetta scores (.json) for all designs
prg_nat_matching: contains inputs: params files (.params) and motifs (.pdb,.cst) used for matching, as well as the source structures (.pdb) from which the motifs were derived.
also contains matches obtained (.pdb) and subsequent designs(filtered->enzdes), .resfiles used for design, and rosetta scores (.json) for all designs
SEC_DATA: contains a .csv file with sec data for all 26 experimentally tested CLAIRE-derived designs, as well as five-point mutants, as shown in Figure 3, Supplementary Figures S8, S9, and S16. Column headers indicate the design identifier. Sets of experiments performed on different dates are separated by a blank column, with the first pair of columns thereafter containing the data for the hplc standard for the following set of experiments. See Methods: Analytical Size Exclusion Chromatography.
CD_DATA: Contains a .csv file with CD scan and thermal melt data for all 3 such characterized CLAIRE designs, A1E, A1P, D2P. See Methods: Circular Dichroism Spectroscopy.
NMR_DATA: Contains Bruker TopSpin acquisition data for all 15N HSQC experiments, including apo and holo chemical shift perturbation data for CLAIRE-derived designs and point mutants shown in Figure 3, Supplementary Figures S10, S11, S12, S17, and S18, as well as titration data for design D2P (Supplementary Figure S14). The NMR acquisition data can be opened and analyzed using Bruker's TopSpin NMR processing software with a freely obtainable academic license. Alternatively, the open source software NMRium (https://app.nmrium.com/)(https://github.com/cheminfo/nmrium) may be used to view and manipulate the data.
The data folders themselves are organized as follows. First, there is a directory labeled by a number (e.g., 1, 2, 3, etc.) which corresponds to the experiment number under a given set of acquisition parameters. This folder contains the raw data, including a 'ser' file containing the raw FIDs, and the 'acqus' and 'acqu2s' files containing the acquisition parameters for the 1st and 2nd dimensions, respectively. The subdirectory 'pdata' contains the processed data, including the file '2rr' containing the processed 2d spectra (corresponding to that shown in relevant figures), as well as the 'procs' and 'proc2s' files containing processing parameters for the 1st and 2nd dimensions, respectively.
The 'audita' and 'auditp' files contain audit trails of the raw and processed data, respectively. These can be checked within the Topspin software via 'show/verify audit trails' to ensure no illegal manipulation of data.
Descriptions of additional files within these data directories can be found in the Topspin user manual.
The subdirectory D2P_2D_Titration contains the N15 HSQC titration data for design D2P with progesterone (Figure S14). Further subdirectories therein (1-6) contain acquisition data for each of the six titration points in the series, corresponding to progesterone concentrations of 0,4.95,14.5,28.3,53.5, and 96.7 micromolar, and protein concentrations of 72,71.3,69.9,67.9,64.3, and 58.1 micromolar, respectively.
The subdirectory N15_HSQC_CSP contains the chemical shift perturbation data (N15 HSQC data with and without target ligand) shown in Figure 3, Figure S10, Figure S11, Figure S12, Figure S17, and Figure S18. Subdirectories indicate the design identifier. Further subdirectories are named according to the experimental conditions, being just the design name for apo data (e.g., 'A1E') or the design name followed by the ligand 3-letter code for holo data (e.g., 'A1E_ESL', 'A1E_PRG' for estriol and progesterone, respectively).
Code/software
The NMR acquisition data can be opened and analyzed using Bruker's TopSpin NMR processing software with a freely obtainable academic license. Alternatively, the open source software NMRium (https://app.nmrium.com/)(https://github.com/cheminfo/nmrium) may be used to view and manipulate the data.
