Small molecule properties define partitioning into biomolecular condensates
Data files
Jul 19, 2024 version files 181.44 GB
-
05232022_Fluorescent_Drug_PC_Condensates.zip
16.66 GB
-
cGASDNA.zip
16.59 GB
-
Condensate_PC_Data_Excel_sheet_File.zip
5.93 MB
-
Control_Buffer_SingleProtein_component_new.zip
12.96 GB
-
Dhh1.zip
8.93 GB
-
FDA-Small_Molecule_Library_Information.zip
3.32 MB
-
ITC_Data.zip
11 MB
-
README.md
16.92 KB
-
Remeasured_Fluorescence_Drug_Partitioning.zip
345.84 MB
-
SH3PRM.zip
14.66 GB
-
SUMOSIM_Extract_MS_1.zip
14.76 GB
-
SUMOSIM_Extract_MS_2.zip
31.69 GB
-
SUMOSIM_Lysate_MS_1.zip
15.97 GB
-
SUMOSIM_Lysate_MS_2.zip
16.73 GB
-
SUMOSIM_Lysate_MS_3.zip
18.65 GB
-
SUMOSIM.zip
13.32 GB
-
Targeted_Metabolomics_Data.zip
145.75 MB
Abstract
Biomolecular condensates regulate cellular function by compartmentalizing molecules without a surrounding membrane. Condensate function arises from specific exclusion or enrichment of molecules. Thus, understanding condensate composition is critical to characterizing condensate function. While principles defining macromolecular composition have been described, understanding of small molecule composition remains limited. Here we quantified partitioning of ~1700 biologically relevant small molecules into condensates composed of different macromolecules. Partitioning varied nearly a million-fold across compounds but was correlated among condensates, indicating disparate condensates are physically similar. For one system, enriched compounds did not generally bind macromolecules with high affinity under conditions where condensates do not form, suggesting partitioning is not governed by site-specific interactions. Correspondingly, a machine learning model accurately predicts partitioning using only computed physicochemical features of the compounds, chiefly those related to solubility and hydrophobicity. These results suggest that a hydrophobic environment emerges upon condensate formation, driving enrichment and exclusion of small molecules.
https://doi.org/10.5061/dryad.fxpnvx10r
The data set provided here consists of raw data files used for determining the partitioning coefficient (PC) data. This calculation is based on the area under the curve (AUC) analysis of metabolites or FDA-approved small molecule drugs. These small molecules are detected and measured their AUC using either targeted or untargeted metabolomics approaches via mass spectrometry.
The second set of data comprises measurements from Isothermal Calorimetry. These measurements focus on the binding of small molecules to proteins.
The third category of data sets includes microscopy data. Specifically, these data set includes the measurements for the partitioning of fluorescent drugs into biomolecular condensates. The acquisition of these data sets is facilitated through confocal fluorescence microscopy.
Description of the data and file structure
# File name(s):#
Targeted_Metabolomics_Data.zip
This dataset contains information on the partitioning of small molecule (50-1000 Da) metabolites into cGASDNA, Dhh1, SH3PRM, and SUMOSIM condensates. It includes data from all replicates for a total of 200 metabolites.Further details about the metabolites can be found in the materials and methods section of the paper.
## Description of the data and file structure
For the detection and quantification of metabolites partitioned into biomolecular condensates, we used a previously described protocol for profiling metabolites extracted from different biological samples. Briefly, for the relative quantification of metabolites partitioned into biomolecular condensates, we used a quantitative polar metabolomics profiling platform using a 5500 QTRAP hybrid triple quadrupole mass spectrometer (AB Sciex). This platform uses a hydrophilic interaction liquid chromatography at high pH (~9.0) with positive/negative ion switching to analyze approximately 200 metabolites (roughly 300 metabolites Q1/Q3 transitions are present in the profile) from a single 15-min liquid chromatography-mass spectrometry acquisition. This method requires no sample manipulation except metabolites are extracted in organic solvent (in our experiment, metabolite extraction carried out using -80 °C methanol, and final methanol concentration in the MS sample is approximately 80%). We used standard method with positive and negative switching for the detection of 200 metabolites. For each sample, we injected 5-μL volume, and the minimum concentration of metabolites was approximately 100 fmoles.
The identity of each metabolite (retention time and specific Q1/Q3MRM transition spectra) confirmed by injecting individual metabolites into the same Mass spectrometer. Using these parameters, we detected and quantified the presence of each metabolite present in individual samples. We acquired the selected reaction monitoring (SRM) data for all the metabolites. Peaks for each data were integrated to generate a chromatographic peak area and used for quantification across the data set. For the analysis, we have used the retention time of each metabolite, area under the detected peak, and the height of the peak. The criteria we were used for the analysis is that the intensity of the peak should be greater than 1e3 (lower limit of detection) and number of Gaussian peaks are 1. After detecting each metabolite and their area, we have made a ratio table for comparing the area of each metabolite in the test samples and compared with the QC samples. QC samples (10 µl supernatant was collected after methanol precipitation from each reaction and pooled together) ran in triplicates before the experiment, in the middle of the experiment, and at the end of the experiment to verify the detection ability of the Mass Spectrometer (MS). For all this analysis, we used MultiQuant 2.0 from AB/SCIEX. All the peaks correspond to each metabolite were further verified by their selected Q1/Q3 transitions and using their elution time. Elution time for each metabolite verify from individual injection of each metabolite under identical LC-MS running conditions. We used relative quantification of metabolites across all the samples, whereas multiple metabolites mixture served as a quality control/reference control.
The files are meticulously organized according to five distinct pool libraries, Withing library, further distinction is made: "A" at the end means droplet fractions, for example S1A corresponds to a droplet fraction, while "B" at the end means the solution fractions, for example S1B corresponds to a solution fraction. Replicates are run next to each other on the same day under identical instrumental conditions. This systematic arrangement facilitates efficient navigation and retrieval of data, enabling researchers to access specific information pertaining to both the pooled libraries and their corresponding fractions with ease.
In all these files HB-80-HILIC-6600-MS means a standard metabolite library. In each of these files, the term 'HB-80-HILIC-6600-MS' refers to a standard metabolite library employed to assess the efficacy of both column chromatography and the mass spectrometer. The abbreviation 'HB' signifies Hamid Baniasadi, the individual overseeing the mass spectrometry facility. Quality Control samples are denoted by 'QC,' while 'start/end' indicates the commencement or conclusion of each set of measurements.
# File names:
(1) Control_Buffer_SingleProtein_Component_new.zip
This dataset encompasses rerun controls designed to assess the degree of precipitation or partitioning of small molecules within buffer or single protein controls. It features data from two replicates for each of the 1500 drug molecules tested across five distinct pooled libraries. Within these libraries, approximately 300 compounds are meticulously pooled from 96-well plates sourced from the Prestwick chemical Libraries.
(2) cGASDNA.zip
This dataset contains information on the partitioning of FDA-approved small molecule drugs into cGAS-45bpDNA condensates. It includes data from two replicates for a total of 1500 drug molecules, data acquired using five different pooled libraries. Each library comprises approximately 300 compounds, which are pooled from 96-well plates sourced from the Prestwick chemical Libraries.
(3) Dhh1.zip
This dataset contains information on the partitioning of FDA-approved small molecule drugs into Dhh1 condensates. It includes data from two replicates for a total of 1500 drug molecules, data acquired using five different pooled libraries. Each library comprises approximately 300 compounds, which are pooled from 96-well plates sourced from the Prestwick chemical Libraries.
(4) SH3PRM.zip
This dataset contains information on the partitioning of FDA-approved small molecule drugs into SH3-PRM condensates. It includes data from two replicates for a total of 1500 drug molecules, data acquired using five different pooled libraries. Each library comprises approximately 300 compounds, which are pooled from 96-well plates sourced from the Prestwick chemical Libraries.
(5) SUMOSIM.zip
This dataset contains information on the partitioning of FDA-approved small molecule drugs into SUMOSIM condensates. It includes data from two replicates for a total of 1500 drug molecules, data acquired using five different pooled libraries. Each library comprises approximately 300 compounds, which are pooled from 96-well plates sourced from the Prestwick chemical Libraries.
(6) SUMOSIM_Lysate_MS_1.zip
(7) SUMOSIM_Lysate_MS_2.zip
(8) SUMOSIM_Lysate_MS_3.zip
(9) SUMOSIM_Extract_MS_1.zip
(10) SUMOSIM_Extract_MS_2.zip
This dataset contains information on the partitioning of FDA-approved small molecule drugs into SUMOSIM condensates prepared in U2OS cell lysate or Xenopus oocyte egg extract. It includes data from two replicates for a total of 1500 drug molecules, data acquired using five different pooled libraries. Each library comprises approximately 300 compounds, which are pooled from 96-well plates sourced from the Prestwick chemical Libraries.
## Description of the data and file structure
For the detection of drug and other small molecules, such as fluorophores, we were used a standard untargeted metabolomic analysis method using a TripleTOF® 6600 System (Sciex). Using this method in positive and negative polarity modes, we were able to detect small molecules based on their retention time and molecular weight and relative quantification of small molecules in any given sample. Samples were prepared using previously described method (See section small molecule partitioning assay), prepared finally for the mass spec in 20% methanol. In this method, we were injected 10-μL volume from each sample into a C18 column with mobile phases of Acetonitrile and water with 0.1 % Formic acid. The minimum concentration of each small molecule injected into mass spectrometer is approximately 200 fmoles. Quality control (QC) samples were prepared by pooling 10 µl of each sample and run in the system before the experiment, in the middle of the experiment, and at the end of the experiment for checking the instrument stability and sample detection. Each individual drug molecule was identified by their characteristic retention time on a C18 column and precursor molecular mass by injecting individual drug molecule libraries in the absence of any biomolecular condensates. Once we confirm the identity and mass accuracy within the threshold limit (± 5 ppm), for the analysis, we used the retention time of each drug molecule, area under the detected peak, and the height of the peak. For all the small molecule analysis, we have used the following criteria: the minimum intensity of the peak should be greater than 1e3, mass accuracy should be within the threshold limit (± 5 ppm), retention time should be within ± 0.2 minutes, and number of Gaussian peaks is 1. After applying these criteria, we calculated area under the curve for each drug molecule and compared with the QC samples. We calculated the coefficient variation (CV) of each small molecules detected in QC samples and eliminated the ones having CV less than 20% for further partitioning coefficient analysis. For calculating the partitioning coefficients (PCs), we simply divided the area under the curve of a given small molecule present in the droplet fraction with corresponding solution fraction.
The files are meticulously organized according to five distinct pool libraries, Withing library, further distinction is made: "A" at the end means droplet fractions, for example S1A corresponds to a droplet fraction, while "B" at the end means the solution fractions, for example S1B corresponds to a solution fraction. Replicates are run next to each other on the same day under identical instrumental conditions. This systematic arrangement facilitates efficient navigation and retrieval of data, enabling researchers to access specific information pertaining to both the pooled libraries and their corresponding fractions with ease. In all these files HB-80-HILIC-6600-MS means a standard metabolite library. In each of these files, the term 'HB-80-HILIC-6600-MS' refers to a standard metabolite library employed to assess the efficacy of both column chromatography and the mass spectrometer. The abbreviation 'HB' signifies Hamid Baniasadi, the individual overseeing the mass spectrometry facility. Quality Control samples are denoted by 'QC,' while 'start/end' indicates the commencement or conclusion of each set of measurements.
# File name(s): ITC_Data.zip
This data file folder contains all the Isothermal Colorimetry (ITC) experiments carried out using SUMO-10R and SIM-10R with small molecules.
## Description of the data and file structure
ITC measurements were performed at 25 °C and 35 °C using a MicroCal PEAQ-ITC calorimeter (Malvern Panalytical). Purified polySUMO (SUMO-10R) and polySIM (SIM-10R) was buffer exchanged and additionally purified by size exclusion chromatography using Superdex200 column with a mobile phase buffer, 25 mM HEPES buffer (pH 7.4) and 150 mM NaCl, at 4 °C, and the fractions were collected, concentrated, and stored in the same buffer. The concentrations of proteins were measured using UV-VIS Spectrophotometer. Proteins were flash frozen in liquid Nitrogen and stored at −80 °C. For the ITC assay, the protein solutions were diluted using the same buffer and prepared a final concentration of 2 µM polySUMO and 2 µM polySIM (20 µM module concentration) in 25 mM HEPES-NaOH (pH 7.4) and 150 mM NaCl (phase separation buffer). For each data point in ITC measurements, 1.9 μL of drug molecule, stock concentration 200 µM, dissolved in the identical phase separation buffer were injected into 0.3 ml of protein in the chamber every 120 s, and 20 injections per experiment. Data for raw ITC and thermodynamic curves, each from one experiment, were downloaded after analysis using Microcal PEAQ-ITC software and plotted using GraphPad Prism.
The details of each compound (drug identifier) provided in the first file folder
# File name(s):
05232022_Fluorescent_Drug_PC_Condensates.zip
Remeasured_Fluorescence_Drug_Partitioning.zip
This data file folder contains all the microscopy images recorded for fluorescent drug molecule partitioning into condensates formed by cGASDNA, Dhh1, SH3PRM and SUMOSIM.
## Description of the data and file structure
Microscopy experiments were carried out in 384-well glass bottom microwell plates (Brooks Life Science Systems: MGB101-1-2-LG-L). For small molecule partitioning experiments using microscopy, scaffold macromolecules were mixed with fluorescent/fluorescently tagged small molecules (100 nM-1 µM, depends on the fluorophore or small molecules) in wells of 384-well plates prepared by using the above-mentioned protocol. Mixtures of scaffold molecules and small molecules were incubated for 1 hour (for cGAS-DNA), 4 hours (for SUMO/SIM), 12 hours (for SH3/PRM), and 20 hours (for Dhh1) at room temperature. Images acquired after incubation using a 60x objective on a Zeiss 780 laser scanning confocal microscope.
# File name(s):
SH3PRM.XSLX
SUMOSIM.XSLX
CGASDNA.XSLX
DHH1.XSLX
This dataset contains information on the partitioning of FDA-approved small molecule drugs into CGASDNA, DHH1, SH3PRM, and SUMOSIM condensates. It includes data from two replicates for a total of 1500 drug molecules, data acquired using five different pooled libraries. Each sheet in the xslx file includes processing of data using various cut offs (such as detection of compounds, mass error, sample to control ratio etc. For more details about the data analysis see the Materials and Method section of the manuscript.)
Each library comprises approximately 300 compounds, which are pooled from 96-well plates sourced from the Prestwick chemical Libraries. Further details about the libraries and molecules can be found in the materials and methods section of the paper.
# File name(s):
FDA-Small_Molecule_Library_Information.zip
This data set includes information about the FDA-approved library obtained from Prestwick and method files used for data analysis.
## Description of the data and file structure
This zip file includes the following files.
Molecules_FDA_Stripped_withoutsalt.sdf
Prestwick_Chemical_Library_Ver20.sdf
Prestwick_Chemical_Library_Ver20.dwar
Prestwick_Chemical_Library_Ver20.txt
FDA_Library_Stripped_Library#1(P1-P4).qmethod
FDA_Library_Stripped_Library#2(P5-P8).qmethod
FDA_Library_Stripped_Library#3(P9-P12).qmethod
FDA_Library_Stripped_Library#4(P13-P16).qmethod
FDA_Library_Stripped_Library#5(P17-P19).qmethod
The files are meticulously organized according to five distinct pool libraries, Withing library, further distinction is made: "A" at the end means droplet fractions, for example S1A corresponds to a droplet fraction, while "B" at the end means the solution fractions, for example S1B corresponds to a solution fraction. Replicates are run next to each other on the same day under identical instrumental conditions. This systematic arrangement facilitates efficient navigation and retrieval of data, enabling researchers to access specific information pertaining to both the pooled libraries and their corresponding fractions with ease.
Code/Software
This data is acquired using Sciex Analyst software, analyzed using Sciex OS software, and finally processed using microsoft excel. The ITC data is acquired using Microcal PEAQ-ITC software, and analyzed using Microcal PEAQ-ITC analysis software. The microscopy data is acquired using Leica LASX Software, and analyzed using FIJI/Image J software. All the .sdf files are analyzed using Datawarrior software. All the .qmethod files are generated and analyzed using SciexOS software.
- Ambadi Thody, Sabareesan; Clements, Hanna D.; Baniasadi, Hamid et al. (2024). Small-molecule properties define partitioning into biomolecular condensates. Nature Chemistry. https://doi.org/10.1038/s41557-024-01630-w
- Thody, Sabareesan Ambadi; Clements, Hanna D.; Baniasadi, Hamid et al. (2022). Small Molecule Properties Define Partitioning into Biomolecular Condensates [Preprint]. Cold Spring Harbor Laboratory. https://doi.org/10.1101/2022.12.19.521099
