Data from: Diversity at the HYP1 locus in potato cyst nematodes does not result from developmentally-programmed somatic mutations
Data files
Jul 21, 2026 version files 1.55 GB
-
cyst_crushing_block
10.54 KB
-
ERR123954_SNVs_filtered.vcf.gz
1.53 GB
-
Gpal_newton_newton.gff3.gz
3.86 MB
-
graph_assembly_gpHYP1_SNVs_filtered.vcf.gz
383.20 KB
-
graph_assembly_gpHYP1.3d73c94.11fba48.8088a73.smooth.final.gfa
557.88 KB
-
manual_examination_old_R9_gpHYP1.pdf
6.22 MB
-
manual_examination_R10_gpHYP1.pdf
2.07 MB
-
manual_examination_R10_grHYP1.pdf
574.36 KB
-
manual_examination_yeast_HYP1.pdf
1.84 MB
-
README.md
3.05 KB
Abstract
In parasites, genetic diversity is valuable fuel for the coevolutionary arms race with their hosts. However, in parasitic nematode worms, low genetic diversity is often observed. Several other parasites and pathogens (e.g., single-celled eukaryotes) generate genetic diversity in an unusual way: they edit the DNA sequence of an important gene in their own genome, instead of waiting for rare spontaneous mutations to occur. This has never been observed in a plant parasite, but we previously described a major parasitism gene (HYP1) in plant-parasitic nematodes with such diverse DNA sequences that it looks like it could perhaps be edited. Now, we generate better data for a simplified genetic system and show that much of the observed HYP1 diversity was actually cryptic DNA sequencing errors. Although there is real genetic diversity at HYP1, it is best explained not by editing but by fundamental evolutionary forces that operate across the genome.
Dataset DOI: 10.5061/dryad.1c59zw4cd
Description of the data and file structure
This dataset largely consists of analysis files and plots of DNA sequence data from Globodera nematodes and transgenic yeast. Additionally, there is a .sat file for creating (or ordering) a custom aluminium block used for extracting eggs from Globodera nematode cysts. One VCF file was produced using a pre-existing short read dataset (European Nucleotide Archive ERR123954).
Files and variables
File: cyst_crushing_block
Description: Dimensions of an aluminium block for crushing Globodera cysts to extract eggs.
File: graph_assembly_gpHYP1.3d73c94.11fba48.8088a73.smooth.final.gfa
Description: Graph genome assembly of the gpHYP1 locus + surrounding sequences, representing the four most common haploytpes in the studied G. pallida population. Can be viewed or manipulated with programs such as bandage or pggb.
File: graph_assembly_gpHYP1_SNVs_filtered.vcf.gz
Description: SNVs called from an alignment of Cas9-enriched R10 nanopore reads against the graph genome assembly, using the most common path (haplotype) as a reference. Invariant sites were artificially added to the VCF by constructing a mock "all sites" invariant VCF for the reference path and merging it into a variant-only VCF.
File: manual_examination_old_R9_gpHYP1.pdf
Description: Plots for manual examination of the nanopore reads from a previous study (doi.org/10.1016/j.xgen.2024.100580) for the Globodera pallida HYP1 gene
File: manual_examination_R10_gpHYP1.pdf
Description: Plots for manual examination of high-accuracy R10 nanopore reads for the Globodera pallida HYP1 gene
File: manual_examination_R10_grHYP1.pdf
Description: Plots for manual examination of high-accuracy R10 nanopore reads for the Globodera rostochiensis HYP1 gene
File: manual_examination_yeast_HYP1.pdf
Description: Plots for manual examination of nanopore reads from the Globodera rostochiensis HYP1 gene, obtained from transgenic yeast transformed with HYP1 and a selectable marker that selects for mutated HYP1 genes.
File: ERR123954_SNVs_filtered.vcf.gz
Description: SNVs called from a pre-existing short-read dataset for G. pallida (European Nucleotide Archive ERR123954) aligned to a linear reference genome and filtered in preparation for population genetic analysis, with invariant sites included.
File: Gpal_newton_newton.gff3.gz
Description: Gene and exon coordinates used to identify four-fold degenerate sites on the linear G. pallida genome assembly (from BioProject PRJNA702104)
Access information
One VCF file was derived from a pre-existing dataset, European Nucleotide Archive ERR123954
