Genome-targeted enrichment and sequencing of human-infecting Cryptosporidium spp.
Data files
Jul 17, 2026 version files 173.70 MB
-
crypto_data.zip
173.30 MB
-
CryptoCap_100k.Rmd
385.54 KB
-
README.md
18.66 KB
Abstract
Cryptosporidium spp. are parasites that cause severe illness in vulnerable human populations. Obtaining pure and sufficient Cryptosporidium DNA from clinical and environmental samples is a challenging task. Oocysts shed in available fecal samples can be limited in quantity, require purification (biased towards dominant strains), and yield limited DNA (<40 fg/oocyst). Here, we use updated genomic sequences from a broad diversity of Cryptosporidium species that have been found to infect humans (C. cuniculus, C. hominis, C. meleagridis, C. parvum, C. tyzzeri, and C. viatorum) to develop and validate a set of 100,000 RNA baits (CryptoCap_100k) with the aim of enriching Cryptosporidium DNA from varied samples. Compared to unenriched libraries, CryptoCap_100k increases the percentage of reads mapping to target genome sequences, increases the depth and breadth of genome coverage, and facilitates analyses of genetic variants in many samples, while decreasing overall costs.
Dataset DOI: 10.5061/dryad.gtht76j0h
Description of the data and file structure
Supplementary Materials for the article “Genome targeted enrichment and sequencing of human-infecting Cryptosporidium spp.”, providing all necessary information to reproduce the analyses and the majority of the figures.
Files and variables
File: CryptoCap_100k.Rmd
Description: R Markdown file «CryptoCap_100k.Rmd»: This R Markdown file is the supplemental code companion to the crypto_data dataset. That is, this R Markdown file presents all the analyses and runs some of them from the manuscript “Genome-targeted enrichment and sequencing of human-infecting Cryptosporidium spp.” using the input files from the crypto_data.zip folder.
Archive: crypto_data.zip
Description: This compressed folder contains subfolders 2, 4–10, and 12–13. Each of which contains input folders and files for analyses corresponding to the numbered Supplementary Notes, organized in the CryptoCap_100k.Rmd file. For example, Supplementary Note #2 uses the files in the “2” folder within crypto_data.zip as input. Below, we expand on what files are contained within each folder:
Subfolder 2: baits_comparisons.txt
Description: Text document that contains summary statistics for two different bait sets (CryptoCap_75k or CryptoWGE v1 and CryptoCap_100k or CryptoWGE v2), such as mean depth and breadth of genome coverage, and percentage of baits that mapped from the mapping of each bait set to 12 Cryptosporidium reference genome sequences.
Subfolder 4:
- *_ct.txt (24 text files)
Description: 24 text files that, for a given Cryptosporidium reference genome sequence and a given fragment size (short or large), report the number of hits of simulated enriched sequences with CryptoCap_100k across different contigs of the reference genome - 18S_RNA folder
Description: 24 text files that, for a given Cryptosporidium reference genome sequence and a given fragment size (short or large), report the number of hits of simulated enriched sequences with CryptoCap_100k against the 18S rRNA database. - gp60 folder
Description: 24 text files that, for a given Cryptosporidium reference genome sequence and a given fragment size (short or large), report the number of hits of simulated enriched sequences with CryptoCap_100k against the gp60 database.
Subfolder: 5
- Lindgreen_species_ct.txt
Description: text document containing the number of hits obtained from mapping the CryptoCap_100k bait set to the Lindgreen database, the species to which it mapped, the phylum, and the domain. - Lindgreen_mapped_both_modified.txt
Description: text document containing the SAM file summary mapping statistics from mapping the CryptoCap_100k bait set to the Lindgreen database. - Lindgreen_mapped_mapq_20.txt
Description: text document containing the filtered hits from the SAM file from mapping the CryptoCap_100k bait set to the Lindgreen database with values of mapq of 20 or larger; it also includes the species to which it mapped, the phylum, and the domain.
Subfolder: 6
- bos_taurus_mapped_modified.txt
Description: text document containing the SAM file summary mapping statistics obtained from mapping the CryptoCap_100k bait set to the Bos taurus reference genome sequence, the species to which it mapped, the phylum, and the domain.
Subfolder: 7
- *_ct.txt (4 text files)
Description: 4 text files that, for a given simulated dataset (three, five, ten, or twelve combined Cryptosporidium species), using a large fragment size, report the number of hits across different contigs of the Crypt10GS reference sequences. - 18SrRNA folder
Description: 4 text files that, for a given simulated dataset (three, five, ten, or twelve combined Cryptosporidium species), using a large fragment size, report the number of hits across different contigs of the 18S rRNA database. - gp60 folder
Description: 4 text files that, for a given simulated dataset (three, five, ten, or twelve combined Cryptosporidium species), using a large fragment size, report the number of hits across different contigs of the gp60 database.
Subfolder: 8
- *_ct.txt (10 text files)
Description: Text documents containing information for the pure-oocyst DNA samples, including the number of hits obtained from mapping sequencing data from unenriched libraries or libraries enriched with CryptoCap_100k to the Crypto10GS database. The corresponding reference genome species is provided in the first column. - pure_oocysts_summary_stats.txt
Description: Text document containing sequencing and mapping summary statistics for pure-oocyst DNA samples from unenriched libraries and libraries enriched with CryptoCap_100k, mapped to their respective reference genome sequences (C. parvum or C. meleagridis).
Subfolder: 9
- *_ct.txt (4 text files)
Description: Text documents containing the number of hits obtained by mapping sequencing data from C. parvum or C. meleagridis DNA samples, prepared as unenriched libraries or libraries enriched with CryptoCap_100k, to the Crypto10GS database. - mixed_infections_NEB-iTru_summary_stats.txt
Description: Text document containing sequencing and mapping summary statistics for C. parvum or C. meleagridis DNA samples from unenriched libraries or libraries enriched with CryptoCap_100k, mapped to their respective reference genome sequences or to the Crypto10GS database. - 18SrRNA folder
Description: 4 text files reporting the number of hits across the 18S rRNA database for each DNA sample. - gp60 folder
Description: 4 text files reporting the number of hits across the gp60 database for each DNA sample.
Subfolder: 10
- C_parvum_bait_dilution_NEB.txt
Description: Text document containing sequencing information and mapping statistics against the C. parvum reference genome sequence and the Crypto10GS database for serially diluted C. parvum DNA samples prepared using the NEB-iTru protocol, with varying volumes of CryptoCap_100k enrichment or no enrichment. - C_par_dilution_iNextEra.txt
Description: Text document containing sequencing information and mapping statistics against the C. parvum reference genome sequence and the Crypto10GS database for serially diluted C. parvum DNA samples prepared using the iNextera-iNext protocol, with varying volumes of CryptoCap_100k enrichment or no enrichment. - iTru_species_renamed folder
Description: 42 text files reporting the number of hits across species from mapping against the Crypto10GS database for each library prepared using the NEB-iTru protocol. - iNextEra_species_renamed folder
Description: 34 text files reporting the number of hits across species from mapping against the Crypto10GS database for each library prepared using the iNextera-iNext protocol.
Subfolder: 12
- seven_clinical_samples_summary_stats.txt
Description: Text document containing sequencing information and mapping of sequencing read statistics against Crypto10GS for seven clinical samples, with libraries enriched for CryptoCap_100K or left unenriched. - *_species.txt
Description: Text files representing each library from a given DNA sample prepared from the set of seven clinical samples, either enriched for CryptoCap_100K or left unenriched. These files report the number of hits across species from the mapping of sequencing reads against the Crypto10GS database. - 18S_rRNA folder
Description: Three text files, each representing a library from a given DNA sample prepared from the set of seven clinical samples, either enriched for CryptoCap_100K or left unenriched, reporting the number of hits across species from mapping of sequencing reads against the 18S rRNA database. - clinical_samples_iTru_summary_stats.txt
Description: Text document containing sequencing information and mapping of sequencing read statistics against Crypto10GS and the C. parvum reference genome sequence for one hundred clinical samples, with libraries enriched for CryptoCap_100K or left unenriched, prepared using the NEB-iTru protocol. - iTru-NEB/species_edited folder
Description: 200 text files for one hundred clinical samples, with libraries enriched for CryptoCap_100K (100) and left unenriched (100), prepared using the NEB-iTru protocol. These files report the number of hits across species from the mapping of sequencing reads against the Crypto10GS database. - clinical_samples_iNext_summary_stats.txt
Description: Text document containing sequencing and mapping information against Crypto10GS and the C. parvum reference genome sequence for one hundred clinical samples, with libraries enriched for CryptoCap_100K or left unenriched, prepared using the iNextEra-iNext protocol. - iNextEra/species_edited folder
Description: 200 text files for one hundred clinical samples, with libraries single-enriched for CryptoCap_100K (100) and left unenriched (100), prepared using the iNextEra-iNext protocol. These files report the number of hits across species from the mapping of sequencing reads against the Crypto10GS database. - iNextEra/double_species_edited folder
Description: 200 text files for one hundred clinical samples, with libraries double-enriched for CryptoCap_100K (100) and left unenriched (100), prepared using the iNextEra-iNext protocol. These files report the number of hits across species from the mapping of sequencing reads against the Crypto10GS database. - clinical_samples_iTruvsiNext_summary_stats.txt
Description: Text document containing comparative information on the sequencing and mapping of sequencing reads against Crypto10GS and the C. parvum reference genome sequence for one hundred clinical samples, with both kinds of libraries, single-enriched with CryptoCap_100K using the iNextEra-iNext protocol or using the NEB-iTru protocol.
Subfolder: 13
- filtered_snps_gatk_all_D.g.vcf
Description: Resulting VCF generated by GATK GenotypeGVCF for libraries produced for one hundred clinical samples that were double-enriched with CryptoCap_100K using the iNextEra-iNext protocol. - filtered_snps_gatk_all_S.g.vcf
Description: Resulting VCF generated by GATK GenotypeGVCF for libraries produced for one hundred clinical samples that were single-enriched with CryptoCap_100K using the iNextEra-iNext protocol. - filtered_snps_gatk_all_U.g.vcf
Description: Resulting VCF generated by GATK GenotypeGVCF for libraries produced for one hundred clinical samples that were left unenriched using the iNextEra-iNext protocol. - vcf_filtered_samples_double.vcf
Description: Filtered VCF containing only SNPs with coverage (DP) between 5 and 200, that were biallelic in the dataset, with less than 40% missing data per SNP, and with samples containing >80% missing data removed. This VCF is for libraries produced for clinical samples that were double-enriched with CryptoCap_100K using the iNextEra-iNext protocol. - vcf_filtered_samples_single.vcf
Description: Filtered VCF containing only SNPs with coverage (DP) between 5 and 200, that were biallelic in the dataset, with less than 40% missing data per SNP, and with samples containing >80% missing data removed. This VCF is for libraries produced for clinical samples that were single-enriched with CryptoCap_100K using the iNextEra-iNext protocol. - vcf_biallelic_unenriched.vcf
Description: Filtered VCF containing only biallelic SNPs in the dataset. This VCF is for libraries produced for clinical samples that were left unenriched using the iNextEra-iNext protocol. - SplitsTree/*-phylip folders
Description: Three folders containing the Phylip file converted from the respective VCF-filtered file and the network files for the corresponding SplitsTree network. These folders represent the corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - Structure/*_results folders
Description: Three folders containing the Structure software output files from the respective corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - Structure/C_hominis folder
Description: Structure software input files (.structure, params folders with main params and extra params, and executable .sh files) and output files from the corresponding libraries produced for UKH clinical samples typed as C. hominis, either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - Structure/C_parvum folder
Description: Structure software input files (.structure, params folders with main params and extra params, and executable .sh files) and output files from the corresponding libraries produced for UKP clinical samples typed as C. parvum, either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - DEploid/*.vcf (3 VCF files)
Description: Three DEploid VCF filtered files generated by VCFTools from the corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - DEploid/testing_haploid/*.vcf (3 VCF files)
Description: Three edited VCF files containing VQSLOD and two values for the AD field, from the corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - DEploid/testing_haploid/*_individual_vcfs folders
Description: Three folders containing the filtered and edited VCF files for each sample from the corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - DEploid/testing_haploid/plaf_file_*.csv (3 CSV files)
Description: Three PLAF files containing the allele frequencies required to run DEploid, from the corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - DEploid/testing_haploid/DEploid_outputs folder
Description: DEploid software output files from the corresponding libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. - nucleotide_diversity/*.recode.vcf
Description: Filtered and edited VCF files generated after running GATK with the diploid option active for subsets of samples (all samples, UKP C. parvum samples only, UKH C. hominis samples only, and UKH C. hominis samples excluding UKH149) from libraries produced for clinical samples that were either double-enriched or single-enriched with CryptoCap_100K, all using the iNextEra-iNext protocol. - nucleotide_diversity/*.sites.pi
Description: Output files from VCFTools reporting per-site nucleotide diversity (.sites.pi) for UKP C. parvum samples, UKH C. hominis samples, and UKH C. hominis samples excluding UKH149, from libraries produced for clinical samples that were either double-enriched or single-enriched with CryptoCap_100K, all using the iNextEra-iNext protocol. - Heterozygosity/*.csv (3 CSV files)
Description: Strata files for libraries produced for clinical samples that were either double-enriched, single-enriched, or left unenriched, all using the iNextEra-iNext protocol. These files contain library type, sample label, and species detected for each sample and were used to support analyses of within-sample variation.
Sharing/Access Information
Other publicly accessible locations of data associated with “Genome-targeted enrichment and sequencing of human-infecting Cryptosporidium spp”:
- https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1061798/
- https://doi.org/10.6084/m9.figshare.29621024
- https://BioRender.com/nfud7lk
- https://BioRender.com/d59hynt
- https://BioRender.com/uqm7uxs
- https://BioRender.com/yysbjqh
- https://BioRender.com/mrcat8s
Code/software
R Studio and the following packages:
· ade4 1.7–23
· adegenet 2.1.11
· ape 5.8–1
· base 4.4.2
· broom 1.0.7
· dartR 2.9.7
· dartR.data 1.0.8
· datasets 4.4.2
· dials 1.3.0
· dplyr 1.1.4
· emmeans 1.10.7
· extrafont 1.0
· forcats 1.0.0
· formatR 1.14
· gdsfmt 1.42.1
· ggplot2 3.5.2
· ggridges 0.5.6
· ggtree 3.14.0
· ggtreeExtra 1.19.0
· graphics 4.4.2
· grDevices 4.4.2
· hierfstat 0.5–11
· infer 1.0.7
· knitr 1.49
· lattice 0.22–6
· lsmeans 2.30–0
· lubridate 1.9.4
· methods 4.4.2
· modeldata 1.4.0
· parsnip 1.2.1
· patchwork 1.3.1
· permute 0.9–7
· plotly 4.10.4
· PopGenReport 3.1
· poppr 2.9.6
· purr 1.0.4
· qqman 0.1.9
· readr 2.1.5
· recipes 1.1.0
· reshape2 1.4.4
· rsample 1.2.1
· scales 1.4.0
· SNPRelate 1.40.0
· starmie 0.1.2
· stats 4.4.2
· stringr 1.5.1
· tibble 3.3.0
· tidymodels 1.2.0
· tidyr 1.3.1
· tidyverse 2.0.0
· tune 1.2.1
· utils 4.4.2
· vcfR 1.15.0
· vegan 2.6–10
· workflows 1.1.4
· workflowsets 1.1.0
· yardstick 1.3.2
