Data from: Engineered reactivity of a bacterial E1-like enzyme enables ATP-driven modification of protein C termini
Data files
Aug 12, 2025 version files 22.22 GB
-
Figure_5D_tiffs.zip
22.23 MB
-
HsLACE_peptide_screen.zip
1.12 GB
-
MccA-A6X_TOF_MS_data.zip
2.51 GB
-
MccA-G4X_TOF_MS_data.zip
2.61 GB
-
MccA-M1X_TOF_MS_data.zip
3.04 GB
-
MccA-N5X_TOF_MS_data.zip
2.48 GB
-
MccA-N7G_nucleophile_screen.zip
1.59 GB
-
MccA-N7X_TOF_MS_data.zip
3.14 GB
-
MccA-R2X_TOF_MS_data.zip
2.68 GB
-
MccA-T3X_TOF_MS_data.zip
2.70 GB
-
pBH4-cfGFP-HpTeCH.gb
12.05 KB
-
pBH4-cfGFP-LjTeCH.gb
12.06 KB
-
pBH4-EcMccB.gb
11.30 KB
-
pBH4-GFP-LACE_D173_internal.gb
10.94 KB
-
pBH4-His-SUMO-GFP-LPETGG.gb
12.38 KB
-
pBH4-His-Tev-GFP-LPETGG.gb
12.47 KB
-
pBH4-His-Tev-Ubc9.gb
10.30 KB
-
pBH4-HpMccB.gb
11.66 KB
-
pBH4-HsMccB.gb
11.34 KB
-
pBH4-LjMccB.gb
11.69 KB
-
pBH4-MBP-TeCH.gb
11.86 KB
-
pBH4-PTP1B1-321-TeCH.gb
11.47 KB
-
pBH4-zEGFR-TeCH.gb
10.08 KB
-
pET28a-eSrtA.gb
8.98 KB
-
pET28a-protein_L-TeCH.gb
10.30 KB
-
Plotted_spectra_for_MccA-A6X_variants.pdf
11.93 KB
-
Plotted_spectra_for_MccA-G4X_variants.pdf
493.91 KB
-
Plotted_spectra_for_MccA-M1X_variants.pdf
448.54 KB
-
Plotted_spectra_for_MccA-N5X_variants.pdf
520.10 KB
-
Plotted_spectra_for_MccA-N7X_variants.pdf
460.63 KB
-
Plotted_spectra_for_MccA-R2X_variants.pdf
15.38 KB
-
Plotted_spectra_for_MccA-T3X_variants.pdf
14.07 KB
-
pPSL937-anti-GFP_rAB-HC-TeCH.gb
22.84 KB
-
Protein_QC_csv_files.zip
42.65 MB
-
README.md
16.77 KB
-
Source_Data_for_Figures.zip
90.48 MB
-
TOF_data_protein_conjugates.zip
199.06 MB
Abstract
In biological systems, ATP provides an energetic driving force for peptide bond formation, but protein chemists lack tools that emulate this strategy. Inspired by the eukaryotic ubiquitination cascade, we developed an ATP-driven platform for C-terminal activation and peptide ligation based on E. coli MccB, a bacterial ancestor of ubiquitin-activating (E1) enzymes that natively catalyzes C-terminal phosphoramidate bond formation. We show that McCB can act on non-native substrates to generate an O-AMPylated electrophile that can react with exogenous nucleophiles to form diverse C-terminal functional groups, including thioesters, a versatile class of biological intermediates that have been exploited for protein semisynthesis. To direct this activity towards specific proteins of interest, we developed the Thioesterification C-terminal Handle (TeCH)-tag, a sequence that enables high-yield, ATP-driven protein bioconjugation via a thioester intermediate. By mining the natural diversity of the MccB family, we developed two additional MccB/TeCH-tag pairs that are mutually orthogonal to each other and to the E. coli system, facilitating the synthesis of more complex bioconjugates. Our method mimics the chemical logic of peptide bond synthesis that is widespread in biology for high-yield in vitro manipulation of protein structure with molecular precision.
https://doi.org/10.5061/dryad.c59zw3rkb
Description of the data and file structure
This dataset contains files related to the use of MccB as an enzymatic tool for C-terminal bioconjugation. Data include plasmid maps for expression of E. coli MccB and homologs; plasmid maps for expression of fusion proteins that can be modified with MccB; raw LC-TOF MS data for modification of peptides and intact proteins in Agilent's .d format; plotted spectra for peptide variants treated with MccB in pdf format; and source data for figures in the associated publication in .csv format.
Plasmid maps
Construction and use of these plasmids is described in Methods.
pBH4-MBP-TeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of maltose binding protein (MBP) fused to a C-terminal TeCH-tag (MRTGNAG)
pBH4-HpMccB.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of Helicobacter pylori MccB.
pBH4-LjMccB.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of Lactobacillus johnsonii MccB.
pBH4-HsMccB.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of Histophilus somni MccB.
pBH4-GFP-LACE_D173_internal.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of GFP with an internal LACE tag following D173.
pBH4-His-SUMO-GFP-LPETGG.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of GFP with an N-terminal SUMO tag and a C-terminal sortase recognition motif (LPETGG).
pBH4-cfGFP-HpTeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of a cysteine-free variant of GFP fused with a C-terminal H. pylori TeCH-tag (MKLSYRG).
pBH4-His-Tev-Ubc9.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of Ubc9, an enzyme used for LACE tag bioconjugation.
pET28a-eSrtA.gb: Plasmid map in GenBank format. This plasmid is derived from Addgene #75144 and encodes a construct for E. coli expression of the engineered sortase A variant eSrtA.
pBH4-EcMccB.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of E. coli MccB.
pBH4-His-Tev-GFP-LPETGG.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of GFP with a C-terminal sortase recognition motif (LPETGG).
pBH4-cfGFP-LjTeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of a cysteine-free variant of GFP fused with a C-terminal L. johnsonii TeCH-tag (MHRIMKG).
pBH4-PTP1B1-321-TeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of the catalytic domain of human PTP1B (residues 1-321) fused to a C-terminal E. coli TeCH-tag (MRTGNAG).
pET28a-protein_L-TeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression for protein L fused to a C-terminal E. coli TeCH-tag (MRTGNAG).
pPSL937-anti-GFP_rAB-HC-TeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of an anti-GFP recombinant antibody with an E. coli TeCH-tag (MRTGNAG) fused to the C terminus of the heavy chain.
pBH4-zEGFR-TeCH.gb: Plasmid map in GenBank format. This plasmid encodes a construct for E. coli expression of an EGFR-targeting affibody with an E. coli TeCH-tag (MRTGNAG) fused to the C terminus.
LC-TOF MS Raw Data
Positional scanning peptide library data, relevant to Figure 4a
Description: The files listed below contain raw data from LC-TOF MS experiments in which a positional scanning peptide library of E. coli MccA (MRTGNAN) (140 peptides in total) was treated with E. coli MccB and ATP. Details are described in Methods. Files are organized by peptide position: for example MccA-M1X corresponds to the panel of peptides in which methionine 1 was substituted with every other canonical amino acid. Each zip file contains three replicates each of twenty experiments corresponding to the indicated position containing each proteinogenic amino acid in Agilent .d format.
Naming: Within the zip file corresponding to each position, the naming convention lists the peptide variant followed by the replicate number. No enzyme controls list the peptide variant, followed by 'noEnz' followed by the replicate number. For example, M1A1.d is replicate 1 of a reaction that contained MccB, ATP, and MccA-M1A, while M1CnoEnz2. d is replicate 2 of a reaction that contained ATP and MccA-M1C, but no MccB.
These data were used to generate the heatmap in Figure 4a.
MccA-M1X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-M1X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
MccA-R2X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-R2X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
MccA-T3X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-T3X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
MccA-G4X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-G4X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
MccA-N5X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-N5X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
MccA-A6X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-A6X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
MccA-N7X_TOF_MS_data.zip: TOF MS data files in Agilent .d format for experiments in which MccA-N7X peptides were treated with MccB and ATP and formation of MccA-NAMP was analyzed.
Nucleophile screen data, relevant to Figure 2e
MccA-N7G_nucleophile_screen.zip: This zip file contains three replicates each of experiments in which MccA-N7G was treated with MccB, ATP, and each of the nucleophiles shown in Figure 2e.
Naming: Within the zip file, the individual files are named according to the nucleophile that was screened, followed by either 'full' for reactions that contained MccB or 'noE' for no-enzyme controls, followed by the replicate number. For example 'allylamine full 2.d' corresponds to replicate 2 of a reaction that contained MccB, ATP, and allylamine.
HsMccB peptide screen, relevant to Figure 6d
HsLACE_peptide_screen.zip: TOF MS data files in Agilent .d format for experiments in which H. somni MccA-N7G and related variants were treated with H. somni MccB, ATP, and a thiol (Mesna or AcCysNHMe) and formation of thiol-modified peptide was analyzed.
Naming: Files are named as 'Hs_', peptide sequence, enzyme concentration, 'MccB', nucleophile identity. For example, Hs_MLGLRGG_10uM_MccB_Mesna.d corresponds to an experiment in which the peptide MLGLRGG was incubated with 10 uM HsMccB, ATP, and Mesna nucleophile.
TOF data in CSV format
Protein QC data, relevant to all figures
Protein_QC_csv_files.zip: This zip file contains TOF data exported to csv for total ion chromatograms, spectra, and deconvoluted spectra for all proteins used in the associated study.
Naming: each file is named with the protein name followed by the type of data in the file. For example, eSrtA TIC.csv contains the total ion chromatogram, eSrtA spec. csv contains the mass spectrum for eSrtA, and eSrtA decon.csv contains the deconvoluted mass spectrum.
Intact protein bioconjugate data, relevant to Figures 3, 4d, 4e, 4f, 5b, 5c, 5e, 6e
TOF_data_protein_conjugates.zip: This zip file contains Agilent .d files for each bioconjugate synthesized in the associated manuscript.
TOF data in PDF format
Plotted_spectra_for_MccA-M1X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-M1X variants were incubated with MccB and ATP, or a no-enzyme control.
Plotted_spectra_for_MccA-R2X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-R2X variants were incubated with MccB and ATP, or a no-enzyme control.
Plotted_spectra_for_MccA-T3X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-T3X variants were incubated with MccB and ATP, or a no-enzyme control.
Plotted_spectra_for_MccA-G4X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-G4X variants were incubated with MccB and ATP, or a no-enzyme control.
Plotted_spectra_for_MccA-N5X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-M1X variants were incubated with MccB and ATP, or a no-enzyme control.
Plotted_spectra_for_MccA-A6X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-M1X variants were incubated with MccB and ATP, or a no-enzyme control.
Plotted_spectra_for_MccA-N7X_variants.pdf: This PDF file shows product and reactant mass spectra for experiments in which MccA-M1X variants were incubated with MccB and ATP, or a no-enzyme control.
TIFF images
TIFF images for Figure 5D
Figure_5D_tiffs.zip: This zip file contains raw TIFF images that were used to make Figure 5D.
Naming: Files are named 'condition', 0000, 'channel'.tiff. For example no_dox_0000_fitc.tiff corresponds to an image of cell that were not treated with dox (no_dox) in the FITC (green) channel.
Source data
Source_Data_for_Figures.zip: This file contains source data for figures 1E, 1F, 2B, 2D, 2E, 3B, 3C, 3D, 3E, 4A, 4B, 4D, 4E, 4F, 5B, 5C, 5D, 5E, 6D, and 6E. This zip file contains the following folders:
Source_Data_for_Fig_1: Contains the following files:
Source Data for Figure 1E - HPLC-MS.xlsx: Contains triplicate datasets for MccB-catalyzed reactions with MccA-N7X variants, where the identity of X is given as the row name. These data were generated using LC-TOF MS to observe the reactant and product. Values under 'full reactant' are peak areas corresponding to the EIC for the reactant; values under 'full product' are peak areas corresponding to the EIC for the AMPylated product. Values under 'noEnz reactant' are peak areas corresponding to the EIC for the reactant in a no-enzyme control; values under 'noEnz product' are peak areas corresponding to the EIC for the AMPylated product in a no-enzyme control.
Source Data for Figure 1E - pyrophosphate release.xlsx: Contains triplicate datasets for MccB-catalyzed reactions with MccA-N7X variants, where the identity of X is given in the name of the tab in the Excel file. These data were generated using an enzyme-coupled assay for pyrophosphate release (EnzChek, ThermoFisher). The first column in each sheet corresponds to the time in minutes; the next three columns are the absorbance at 360 nm that results from release of pyrophosphate due to MccB-catalyzed ATP turnover.
Source Data for Figure 1F.xlsx: This file contains kinetics data for Michaelis-Menten kinetic analysis of MccB activity on MccA variants. Each sheet contains a table with the first column being substrate concentration and the next three columns corresponding to the rate of substrate release at that concentration.
Source_Data_for_Fig_1: Contains the following files:
Source Data for Figure 2B.xlsx: Contains a spreadsheet with ion counts versus retention time on LC-TOF MS for the product of MccB activity on MccA and MccA-N7G; also contains no enzyme controls.
Source Data for Figure 2D.xlsx: A spreadsheet with data for plotting mass spectra. X column is m/z and Y column is ion counts. Experiments are described in the tab name.
Source Data for Figure 2E.xlsx: Source data for making radial heatmap. First column contains nucleophile identity; next three columns are % conversion when MccB is present; next three columns are % conversion when MccB is absent.
Source_Data_for_Fig_3 contains the following folders:
Source Data for Figure 3B: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for an MccB timecourse with GFP-TeCH in the presence of Mesna
Source Data for Figure 3C: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for an MccB-catalyzed reaction in the presence of Mesna with the following TeCH-tagged proteins: anti-GFP rAb; MBP; protein L; PTP1B; and zEGFR
Source Data for Figure 3D: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for an MccB-catalyzed reaction in the presence of Cys with GFP-TeCH; reaction of the Cys conjugate with biotin-maleimide; and reaction of the Cys conjugate with Cy5-maleimide.
Source Data for Figure 3E: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for an MccB-catalyzed reaction in the presence of MesnA with GFP-TeCH in the presence of the peptide CGAGSAz; and the same file for reactions of the CGAGSAz conjugate with biotin-dibenzocyclooctyne
Source_Data_for_Fig_4 contains the following files/folders:
Source Data for Figure 4A.csv: Table in which rows are labeled with amino acids and columns are labeled with peptide positions. Values in cells are % conversion for the MccA (row name) variant at position (column name)
Source Data for Figure 4B.xlsx: Table that contains data use to construct relative activity heatmap for different MccB homologs. The first colum contains the name of the MccB homolog; the next three columns contain triplicate rate measurements for that homolog vs E. coli MccA, H. pylori MccA, and L. johnsonii MccA
Source Data for Figure 4D: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for homolog MccB reactions with cognate GFP-TeCH in the presence of Mesna
Source Data for Figure 4E: Each tab of this Excel file contains an exported deconvoluted mass spectrum corresponding to GFP-TeCH treated with MccB for the MccB/MccA species indicted in the tab.
Source Data for Figure 4F: Contains exported pmod-deconvoluted mass spectra for a mixture of GFP-TeCHs treated with the homolog (Ec, Hp, Lj, or Hs) indicated in the filename.
Source_Data_for_Fig_5 contains the following folders/files:
Source Data for Figure 5B: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for MccB/subtiligase catalyzed reactions with GFP-TeCH in the presence of Mesna and AFAGAGSazK
Source Data for Figure 5C: Contains TOF data exported to csv for total ion chromatograms, spectra and deconvoluted spectra for MccB/subtiligase catalyzed reactions with TeCH-tagged anti-GFP rAb, MBP, protein L, and PTP1B in the presence of Mesna and AFAGAGSazK
Source Data for Figure 5D: Contains raw, uncropped TIFF images for the doxycline treatment of a doxycycline-inducted GFP-TM cell line (plus_dox) and a no doxycycline control (no_dox)
Source Data for Figure 5E: Contains TOF deconvoluted mass spectra for full sortase/MccB/subtiligase reaction (full) and control lacking both enzymes (-enz); lacking MccB (-MccB); lacking subtiligase (-SL); and lacking sortase (-eSrtA) in either one-pot or telescope format.
Source_Data_for_Fig_6 contains the following files/folders:
Source Data for Figure 6D.xlsx: Excel sheet with peptide substrate names in bold. Under each peptide, three different nucleophiles (Mesna, Mesna/TCEP, and AcCysNHMe were tested with the indicate concentrations of HsMccB. Table show reactant area, product area, and % conversion for each MccB concentration.
Source Data for Figure 6E: Contains TOF deconvoluted mass spectra for LACE-tag GFP/Ubc9 reactions with a synthetic peptide thioester (LRLRGG-Mes) or an enzyme synthesize thioester (MLGLRGG-Mes) as well as a no peptide control
Code/software
Plasmids maps in GenBank format can be imported into Benchling, an electronic lab notebook program, and analyzed. They are also compatible with many other softwares for editing plasmid maps (e.g., ApE, SnapGene).
Agilent .d files can be opened using Agilent MassHunter Qualitative Analysis 10.0. Our workflow involved viewing MS data at specific points in the chromatogram and generated extracted ion chromatograms using built-in features of the software.
Access information
Other publicly accessible locations of the data:
- n/a
Data was derived from the following sources:
- n/a
This deposition includes 1) plasmid maps for expression constructs used in the relevant manuscript; and 2) raw data from an Agilent LC-TOF MS instrument in the *.d format. Raw LC-TOF MS data were processed in Agilent MassHunter BioConfirm v10.0 (intact protein data) or Agilent MassHunter Qualitative Analysis 10.0 (peptide data). The data deposited here has not been processed. For experiments described in our manuscript, intact protein spectra were deconvoluted using the Maximum Entropy algorithm. Peptide data were analyzed by extracting spectra or generating extracted ion chromatograms.
