High-throughput discovery of transmembrane helix dimers from human single-pass membrane proteins with TOXGREEN sort-seq
Data files
Aug 19, 2025 version files 94.51 MB
-
allSeqs17.csv
6.28 MB
-
allSeqs19.csv
4.20 MB
-
allSeqs21.csv
6.95 MB
-
example_reference_file.csv
1.04 MB
-
example.fastq
9.36 MB
-
fastqToTxt.pl
7.30 KB
-
O75056_syndecan-3_AF3_model_0.pdb
506.39 KB
-
O75056_syndecan-3_CATM_model_1.pdb
7.68 MB
-
P02786_TfR1_AF3_model_4.pdb
946.58 KB
-
P02786_TfR1_CATM_model_1.pdb
1.14 MB
-
P16410_CTLA4_AF3_model_3.pdb
273.19 KB
-
P16410_CTLA4_CATM_model_1.pdb
539.49 KB
-
P78380_OLR1_AF3_model_2.pdb
343.34 KB
-
P78380_OLR1_CATM_model_1.pdb
70.25 KB
-
Q13591_sema-5A_AF3_model_1.pdb
1.34 MB
-
Q13591_sema-5A_CATM_model_1.pdb
2.80 MB
-
Q6P7N7_TMEM81_AF3_model_0.pdb
316.64 KB
-
Q6P7N7_TMEM81_CATM_model_1.pdb
8.29 MB
-
Q6UW88_Epigen_AF3_model_3.pdb
189.45 KB
-
Q6UW88_Epigen_CATM_model_1.pdb
8.44 MB
-
Q6UXE8_BTNL3_AF3_model_4.pdb
583.18 KB
-
Q6UXE8_BTNL3_CATM_model_3.pdb
3.97 MB
-
Q8N6P7_IL22R1_AF3_model_0.pdb
701.68 KB
-
Q8N6P7_IL22R1_CATM_model_1.pdb
1.18 MB
-
Q8NEA5_C19orf18_AF3_model_0.pdb
267.97 KB
-
Q8NEA5_C19orf18_CATM_model_2.pdb
137.19 KB
-
Q8NFY4_sema-6D_AF3_model_2.pdb
1.33 MB
-
Q8NFY4_sema-6D_CATM_model_1.pdb
1.56 MB
-
Q8TDQ0_HAVCR2_AF3_model_3.pdb
371.15 KB
-
Q8TDQ0_HAVCR2_CATM_model_1.pdb
6.80 MB
-
Q9H3T3_sema-6B_AF3_model_4.pdb
1.06 MB
-
Q9H3T3_sema-6B_CATM_model_1.pdb
6.42 MB
-
Q9NVM1_EVA1B_AF3_model_2.pdb
204.30 KB
-
Q9NVM1_EVA1B_CATM_model_1.pdb
4.47 MB
-
Q9ULG6_CCPG1_AF3_model_0.pdb
969.49 KB
-
Q9ULG6_CCPG1_CATM_model_1.pdb
3.76 MB
-
README.md
10.01 KB
-
SortGateInfo.csv
5.47 KB
Abstract
The oligomerization of the transmembrane helices of single-pass membrane proteins is crucial to biological function, and its misregulation can lead to many diseases. The study of transmembrane helix oligomerization is facilitated by the availability of genetic reporter assays, which are essential tools for understanding the organization and biology of single-pass systems. In particular, reporter assays are crucial for mapping the oligomerization interfaces of transmembrane helices through scanning mutagenesis, but their application is limited by the need to clone and measure each construct individually. Here, we present “TOXGREEN sort-seq”, a high-throughput version of the TOXGREEN assay that enables the direct measurement of transmembrane helix oligomerization in large libraries using fluorescence-activated cell sorting and next-generation sequencing. We show that TOXGREEN sort-seq is robust and reproduce the direct measurements of individual constructs with good accuracy and sensitivity. The method produced high-quality mutational profiles from a library of 17,400 constructs designed to probe the interface of 100 potential GASrightdimers predicted from sequences of human single-pass membrane proteins. We report the validated structural model of twelve dimers involved in a variety of biological functions, including immune response (interleukin-22 receptor subunit alpha-1, butyrophilin-like protein 3, hepatitis A virus cellular receptor 2), transport (transferrin receptor protein 1), and cell-surface signaling and proliferation (syndecan-3; semaphorins 5A, 6B and 6D). Remarkably, all three semaphorins in the dataset formed strong dimers and produced mutational profiles consistent with the computational structure. These findings open the possibility that dimerization may be relevant to these proteins’ activity and provide a validated interface for assessing their biological role.
Dataset DOI: 10.5061/dryad.qnk98sfv8
This repository contains:
- Structural model of transmembrane helix dimers (PDB files)
- Sort-seq Data
- NGS data processing script (with example input data files)
Structural model of transmembrane helix dimers (PDB files)
Description of the data and file structure
The files include the CATM and AlphaFold3 (AF3) models of the twelve GAS-right dimers validated by the mutagenesis. The files also include the CATM and AF3 models of the three proteins whose mutagenesis matches an AF3 model but not a CATM model. The CATM and AF3 models have been aligned with each other based on the transmembrane helices.
Name scheme:
<Uniprot Code><protein name abbreviation><modeling program><model number><.pdb>
Files of proteins matching the CATM model:
O75056_syndecan-3_AF3_model_0.pdbO75056_syndecan-3_CATM_model_1.pdbP02786_TfR1_AF3_model_4.pdbP02786_TfR1_CATM_model_1.pdbQ13591_sema-5A_AF3_model_1.pdbQ13591_sema-5A_CATM_model_1.pdbQ6P7N7_TMEM81_AF3_model_0.pdbQ6P7N7_TMEM81_CATM_model_1.pdbQ6UW88_Epigen_AF3_model_3.pdbQ6UW88_Epigen_CATM_model_1.pdbQ6UXE8_BTNL3_AF3_model_4.pdbQ6UXE8_BTNL3_CATM_model_3.pdbQ8N6P7_IL22R1_AF3_model_0.pdbQ8N6P7_IL22R1_CATM_model_1.pdbQ8NEA5_C19orf18_AF3_model_0.pdbQ8NEA5_C19orf18_CATM_model_2.pdbQ8NFY4_sema-6D_AF3_model_2.pdbQ8NFY4_sema-6D_CATM_model_1.pdbQ8TDQ0_HAVCR2_AF3_model_3.pdbQ8TDQ0_HAVCR2_CATM_model_1.pdbQ9H3T3_sema-6B_AF3_model_4.pdbQ9H3T3_sema-6B_CATM_model_1.pdbQ9NVM1_EVA1B_AF3_model_2.pdbQ9NVM1_EVA1B_CATM_model_1.pdb
Files of proteins matching only the AF3 model:
P16410_CTLA4_AF3_model_3.pdbP16410_CTLA4_CATM_model_1.pdbP78380_OLR1_AF3_model_2.pdbP78380_OLR1_CATM_model_1.pdbQ9ULG6_CCPG1_AF3_model_0.pdbQ9ULG6_CCPG1_CATM_model_1.pdb
Code/software
PDB files are viewable with PyMol or other software that can read Protein Data Bank coordinate files.
Sort-seq Data
This repository contains processed data from high-throughput TOXGREEN sort-seq experiments used in the analysis presented in our manuscript. The data include deep sequencing results from three length libraries of single-pass membrane protein variants (wild-types and point mutants), as well as data describing the GFP fluorescence gates used during fluorescence-activated cell sorting.
The following files contain processed sequencing data for three libraries of different transmembrane domain lengths:
allSeqs17.csvallSeqs19.csvallSeqs21.csv
Each row in these files corresponds to a unique sequence observed in a specific library, replicate, and fluorescence bin.
Column Descriptions:
| Column Name | Description |
|---|---|
TM |
Amino acid sequence of the transmembrane construct |
Count |
Number of reads for this sequence in the corresponding bin |
Percent |
Percentage of total reads in the bin represented by this sequence |
ID |
UniProt ID of the wild-type protein |
Name |
Protein name (as annotated in UniProt) |
StartRes |
Starting residue number of the TM segment in the full-length wild-type protein |
Mutation |
Single-point mutation, formatted as WTPosMut (e.g., A123V) |
MutRes |
Position of the mutation within the TM sequence. N/A represents not applicable as wild-type sequences do not have a mutation. |
Library |
Library transmembrane construct length (17, 19, or 21) |
Replicate |
Experimental replicate number |
Bin |
FACS fluorescence bin from which the sequence was recovered |
Sort Gate Information File
SortGateInfo.csv
This file contains data from the FACS gating and binning used during the sorting of each library replicate.
Column Descriptions:
| Column Name | Description |
|---|---|
Bin |
Fluorescence bin number during sorting |
ReflowMedian |
Median GFP fluorescence of the reflow test sort (used to assess sorting bin quality) |
ReflowMean |
Mean GFP fluorescence of the reflow test sort |
Fraction |
Fraction of the total cell population falling within the bin during reflow |
SetMedian |
Median GFP fluorescence of the final sorted population in the gated bin |
SetMean |
Mean GFP fluorescence of the final sorted population in the gated bin |
SetMin |
Minimum GFP fluorescence obtained in the final sorted population in the gated bin |
SetMax |
Maximum GFP fluorescence obtained in the final sorted population in the gated bin |
Library |
Library length (17, 19, or 21) |
Replicate |
Sorting replicate number |
NGS data processing script
This script fastqToTxt.pl processes NGS data in .fastq format by extracting high-quality reads, translating DNA sequences into protein sequences, and matching them against a reference list. It outputs a count and frequency of each detected sequence and preserves annotation information if the sequence is found in a reference sequence file.
Features
- Reads
.fastqfiles and filters out low-quality reads. - Extracts regions between given forward and reverse primers.
- Translates the DNA region into an amino acid sequence.
- Filters sequences by expected N- and C-terminal residues.
- Matches protein sequences to a reference sequence file.
- Reports the count, percentage, and metadata for matched sequences.
Input Files
1. FASTQ File (--seqFile)
A standard .fastq file with 4-line entries (sequence on line 2, quality on line 4).
2. Reference File (--refFile)
A tab-delimited file with 7 columns, where:
| Column | Description |
|---|---|
| 0 | (optional metadata) |
| 1 | (optional metadata) |
| 2 | (optional metadata) |
| 3 | (optional metadata) |
| 4 | (optional metadata) |
| 5 | (optional metadata) |
| 6 | Protein sequence (used for matching) |
The protein sequence in column 6 is the key used to identify matches, and columns 0–5 are printed as annotation when a match is found.
Usage
perl fastqToTxt.pl --refFile reference.csv --seqFile reads.fastq --direction 1
Arguments
--refFile: Path to the tab-delimited reference file.--seqFile: Path to the.fastqfile to be analyzed.--direction: Direction of sequencing reads:1= forward2= reverse
Output
The script prints to standard output a tab-delimited table containing:
<Protein Sequence> <Count> <Fraction of total good reads> <Reference Annotation>
If the sequence is found in the reference file, the annotation from columns 0–5 is printed. If not, the script checks for special fallback cases or prints Unknown.
Notes
- Only sequences with expected quality (≤1 expected error) are retained.
- Primers are hard-coded and must match the sequences used in library construction.
- Sequences must start with
"AS"and end with"L"after translation and are trimmed accordingly. - Matches are based on exact protein sequences; partial or fuzzy matches are not considered.
- A default count threshold of 10 is applied to reduce noise.
Example input files
Run the script with the example input files provided, example_reference_file.csv and example.fastq with:
perl fastqToTxt.pl --refFile example_reference_file.csv --seqFile example.fastq --direction 1
Notes
- Sequence data files reflect processed reads that passed quality and frequency filters.
- Bin statistics in
SortGateInfo.csvwere used to reconstruct fluorescence values for each construct. - All fluorescence values are reported in arbitrary units based on the GFP signal detected during sorting.
Contact
Alessandro Senes, senes@wisc.edu
Access information
Other publicly accessible locations of the data:
- The CATM files are available at catm.biochem.wisc.edu/CATM
- The AlphaFold files can be computed at the AlphaFold server https://alphafoldserver.com/
Data was derived from the following sources:
- The structural predictions were made from the protein sequence from the Uniprot database: https://www.uniprot.org/
