Data from: Virome-wide ubiquitin ligase discovery reveals diverse mechanisms of immune evasion
Data files
Jul 09, 2026 version files 59.40 GB
-
DataS1_aGFPnb_Raw_DNAseq.zip
15.33 GB
-
DataS2_UbCRISPR_Raw_DNAseq.zip
44.07 GB
-
DataS3_AlphaFold3-requests.zip
166.68 KB
-
README.md
3.12 KB
Abstract
Viruses are intracellular parasites that reprogram the host proteome to promote replication and evade immune recognition. We applied a virome-wide library of ~10,000 open reading frames to discover viral ubiquitin ligases, mapping their mechanisms of degradation and host substrates using targeted CRISPR screens and proteomics. These viral effectors can be classified as canonical ligases that mimic host E3s, hijackers that redirect host E3s, and non-canonical ligases that rewire Cullin-RING ligase machinery. These diverse strategies of virus-mediated degradation converged on immune-related substrates, including JAK1 and CUL1b-TrCP, underscoring immune evasion as a major driver of viral ubiquitin ligase evolution. Our findings reveal new viral strategies for exploiting the ubiquitin–proteasome system with potential for therapeutic targeting. The supplementary data sets provided here correspond to raw sequencing data and structural predictions presented in the work.
https://doi.org/10.5061/dryad.pzgmsbd3d
Data description
Data S1 contains the raw data related to the GFP nanobody screen (Fig. 1F). A library of viral open reading frames was fused to a GFP-specific nanobody and screened for loss of GFP fluorescence intensity relative to a DsRed control. Significance was calculated by comparing barcode abundance in the bottom 1% GFP/DsRed cells relative to input using MAGeCK.
Data S2 contains the raw data related to the ubiquitin-focused CRISPR screens (Fig. 2B, 3B, 4B, 5B, 6B). Cell lines expressing an individual viral gene fused to GFP (Fig. 2B Rotavirus NSP1, Fig. 3B TevPV-M, Fig. 4B. Adana virus NSs, Fig. 5B Razdan virus NSs, Fig. 6B ASFV MGF505-9R) were transduced with a library of guide RNAs targeting genes related to ubiquitination. Cells were sorted based on increased GFP fluorescence intensity relative to DsRed to identify host factors involved in viral protein degradation. Significance was calculated by comparing barcode abundance in the top 5% GFP/DsRed cells relative to input using MAGeCK.
Data S3 contains input files for Host-Virus AlphaFold 3 predictions. Fig. 4F Adana virus NSs in complex with human JAK1 pseudokinase domain (545-855). fig. S2B Rotavirus NSP1 genes in complex with CUL3 and ELOC. fig. S3A TevPV-M in complex with BTBD1 and CUL3 N-terminus (1-418). fig. S4E Adana virus NSs and related genes in complex with JAK1 pseudokinase domain (545-855). fig. S5C Razdan virus NSs in complex with CUL3 N-terminus (1-418). fig. S6A ASFV MGF505-9R and MGF360-10L monomers. File naming convention is fold_dateYYMMDD_gene1_gene2_geneX_job_request.json.
Description of the data and file structure
DataS1_aGFPnb_Raw_DNAseq.zip
gz_files - A folder containing all of the raw DNA sequencing data as fastq.gz files. Replicates A and B represent independent infections, selections, and sorts. 'input' refers to sample prior to sorting. 'lo' refers to sample after sorting.
scripts - A folder containing analysis scripts for processing raw data. trim_align_all.sh is a wrapper script that calls trim_align2.sh for individual .gz sequencing files. Modules used: bowtie/1.2.2, gcc/6.2.0, samtools/1.3.1, python/3.7.4, cutadapt/2.5
EF-18-44.ref_table.fa - Reference fasta for alignment purposes.
DataS2_UbCRISPR_Raw_DNAseq.zip
gz_files - A folder containing all of the raw DNA sequencing data as fastq.gz files, organized into subfolders based on figure
UbCRISPRanalysis.txt - command line instructions for analysis of CRISPR screening data. Modules used: gcc/14.2.0, bowtie/1.2.2, samtools/1.21, python/3.13.1, conda/miniforge3/24.11.3-0
Ublib_CRISPRguides.fasta - Reference fasta for alignment purposes
UbCRISPRsamples.csv - file mapping sequencing data to samples
DataS3_AlphaFold3-requests.zip
AlphaFold3 input files arranged by corresponding figure.
job_request.json - input json file showing starting sequences (can be viewed in text editor and used to run AF3 predictions)
