A nuclear genome assembly of an extinct flightless bird, the little bush moa
Data files
Apr 01, 2024 version files 2.39 GB
-
candidate_genes.tar.gz
9.68 MB
-
draft_genome_assemblies.tar.gz
587.55 MB
-
emu-moa_alignments_mapDamage_corrected_assembly.tar.gz
704.25 MB
-
emu-moa_alignments_original_assembly.tar.gz
707.48 MB
-
MHC_alignments.zip
210.63 KB
-
microsatellites.tar.gz
298.74 KB
-
moa_mtDNA_CR_phylogeny.tar.gz
10.91 KB
-
OR_genes.zip
926.94 KB
-
parse_mapDamage_pileup.pl
3.18 KB
-
pop_gen.zip
156.17 MB
-
README.md
23.90 KB
-
Sackton_et_al._2019_update.tar.gz
216.44 MB
-
sensory_gene_alignments.zip
182.84 KB
-
species_codes.txt
3.29 KB
-
TE_analysis.zip
1.79 MB
Abstract
We present a draft genome of the little bush moa (Anomalopteryx didiformis) - one of approximately nine species of extinct flightless birds from Aotearoa, New Zealand - using ancient DNA recovered from a fossil bone from the South Island. We recover a complete mitochondrial genome at 249.9X depth of coverage and almost 900 Mb of a male moa nuclear genome at ~4-5X coverage, with sequence contiguity sufficient to identify more than 85% of avian universal single-copy orthologs. We describe a diverse landscape of transposable elements and satellite repeats; estimate a long-term effective population size of ~240,000; identify a diverse suite of olfactory receptor genes and an opsin repertoire with sensitivity in the UV range; show that the wingless moa phenotype is likely not attributable to gene loss or pseudogenization; and identify potential function-altering coding sequence variants in moa that could be synthesized for future functional assays. This genomic resource should support further studies of avian evolution and morphological divergence.
https://doi.org/10.5061/dryad.d51c59zxp
Description of the data and file structure
GENERAL INFORMATION
1. Title of Dataset: A nuclear genome assembly of an extinct flightless bird, the little bush moa
2. Author Information
A. Principal Investigator Contact Information
Name: Scott V. Edwards
Institution: Harvard University
Address: Dept. of Organismic and Evolutionary Biology, Harvard University, 26 Oxford Street, Cambridge MA, 02138 USA
Email: sedwards@fas.harvard.edu
3. Date of data collection (single date, range, approximate date): 2000-2018
4. Geographic location of data collection: South Island, New Zealand
5. Information about funding sources that supported the collection of the data: Natural Science and Engineering Research Council of Canada (NSERC, to A.J.B.); Royal Ontario Museum Governors Fund (to A.J.B.); National Science Foundation (NSF grant DEB 1355343 (EAR 1355292) to A.J.B. and S.V.E.); Japan Society for the Promotion of Science (KAKENHI Grant Number JP20K06767 to K.K.J.)
DATA & FILE OVERVIEW
1. File List:
A) candidate_genes.tar.gz
B) draft_genome_assemblies.tar.gz
C) emu-moa_alignments_original_assembly.tar.gz
D) emu-moa_alignments_mapDamage_corrected_assembly.tar.gz
E) MHC_alignments.zip
F) microsatellites.tar.gz
G) moa_mtDNA_CR_phylogeny.tar.gz
H) OR_genes.zip
I) parse_mapDamage_pileup.pl
J) pop_gen.zip
K) Sackton_et_al._2019_update.tar.gz
L) sensory_gene_alignments.zip
M) species_codes.txt
N) TE_analysis.zip
2. Relationship between files, if important: None
3. Additional related data collected that was not included in the current data package: None
4. Are there multiple versions of the dataset? No
A. If yes, name of file(s) that was updated: NA
i. Why was the file updated? NA
ii. When was the file updated? NA
#########################################################################
DATA-SPECIFIC INFORMATION FOR: candidate_genes.tar.gz
Archive directory structure:
candidate_genes
├── hyphy_species_tree.tre
├── mapDamage_corrected_assembly
│ ├── DCHS1
│ ├── DVL1
│ ├── DYNC2H1
│ ├── EVC
│ ├── FAT1
│ ├── FGF10
│ ├── FGF8
│ ├── GLI2
│ ├── GLI3
│ ├── HOXA
│ ├── HOXD
│ ├── IFT122
│ ├── KIF7
│ ├── OFD1
│ ├── SALL4
│ ├── SHH
│ ├── TALPID3
│ ├── TBX5
│ ├── WDR34
│ └── WNT2B
└── original_assembly
├── DCHS1
├── DVL1
├── DYNC2H1
├── EVC
├── FAT1
├── FGF10
├── FGF8
├── GLI2
├── GLI3
├── HOXA
├── HOXD
├── IFT122
├── KIF7
├── OFD1
├── SALL4
├── SHH
├── TALPID3
├── TBX5
├── WDR34
└── WNT2B
Description:
File hyphy_species_tree.tre provides the input topology in Newick format that was used for HyPhy analyses.
Higher-level relationships in birds followed Jarvis et al. (2014; Science 346:1320–1331),
with paleognath relationships following Cloutier et al. (2019; Syst. Biol. 68:937-955),
galliforms following Wang et al. (2013; PLoS One 8:e64312), cormorants following Burga et al.
(2017; Science 356:eaal3345), and passerines following Barker et al. (2004; Proc Natl Acad Sci
U S A 101:11040–11045) and Harshman et al. (2007; Classification and phylogeny of birds.
In: Jamieson BGM, ed. Reproductive biology and phylogeny of birds. Enfield, NH: Science Publishers,
Inc. p. 1–35). The full topology was pruned to match the available sequence data for each gene with
ETE v. 3 (Huerta-Cepas et al. 2016; Mol Biol Evol. 33:1635–1638).
Subdirectory mapDamage_corrected_assembly/ includes results of analyses using the mapDamage corrected version
of the genome assembly as input, and subdirectory original_assembly/ includes results using the original
(uncorrected) version of the moa genome assembly.
Within these top-level subdirectories, separate subdirectories are included for each candidate gene.
(note: HOXA/HOXD individual genes are grouped into common subdirectories)
Within each gene-level subdirectory is:
1\) a tab-delimited text file with accessions and additional information for all sequences
retrieved from GenBank
2\) the coding sequence nucleotide alignment and amino acid translations for all species
in Phylip format
3\) an alignment of the moa translated protein sequence to the ancestral reconstruction of
the common ancestor of moa and tinamous used for PROVEAN analysis, in fasta format
(for original moa assembly only)
4\) a subdirectory containing manually curated gene models for the 10 paleognath genomes from
Sackton et al. (2019, Science 364: 74-78), and for the little bush moa, each in gene
feature format (GFF3). GFF coordinates are given relative to annotated gene start/end;
the corresponding genomic coordinates are provided for all taxa in the included BED file.
#########################################################################
DATA-SPECIFIC INFORMATION FOR: draft_genome_assemblies.tar.gz
Archive directory structure:
draft_genome_assemblies
├── mapDamage_corrected_assembly
│ ├── anoDid_nucDNA_mapDamage.fasta.gz
│ └── anoDid_nucDNA_mapDamage_NCBI_mask.bed
└── original_assembly
├── anoDid_nucDNA.fasta.gz
└── anoDid_nucDNA_NCBI_mask.bed
Description:
The draft genome assemblies used for all analyses had some regions masked during
processing for NCBI submission. These regions largely represent untrimmed adapter
sequence that was incorporated into the genome assemblies. In this directory archive,
we provide the draft genome assemblies prior to contaminant screening for both the original
assembly and mapDamage corrected version of the nuclear genome assembly, and BED format files
listing the coordinates that were masked during NCBI contaminant screening. Masking these regions
in the provided draft genome assemblies will yield sequences that are identical to the final versions
released on NCBI. Note that all genome statistics and datasets described in the moa genome manuscript
associated with this data release report results for the final NCBI genomes. However, a few
loci in already published manuscripts (Sackton et al. 2019; Science 364:74-78 and Cloutier
et al. 2019; Syst. Biol. 68:937-955) will be affected by the NCBI contaminant masking.
#########################################################################
DATA-SPECIFIC INFORMATION FOR: emu-moa_alignments_original_assembly.tar.gz
Archive directory structure:
emu-moa_alignments_original_assembly
├── droNov-anoDid_alignments
├── example_coordinates.bed
└── map_droNov-anoDid_coordinates.pl
Description:
Subdirectory droNov-anoDid_alignments/ contains a file in FASTA format for each scaffold in the original
(i.e. mapDamage uncorrected) nuclear genome assembly for little bush moa (anoDid) with moa sequence aligned to the
corresponding scaffold in the reference emu genome (droNov, corresponding to GenBank accession GCA_003342905.1).
Note: emu scaffold names were retained for moa, e.g. scaffold_100 in emu is also called scaffold_100 in moa.
The moa sequences used here correspond to the final genome assemblies released on NCBI (i.e. with regions identified
as contaminant during NCBI processing masked).
Note: genome coordinates are not always identical between the original and mapDamage corrected moa assemblies due to
removal of leading/trailing Ns from the mapDamage assembly scaffolds. Users wishing to output both versions of the moa
sequence should therefore run the provided script (see below) for each provided data set (i.e. both the original and
mapDamage corrected assembly versions).
Because of the fragmented nature (low contig N50) of the moa nuclear genome assembly, users might
find that using coordinates from the reference emu genome to retrieve moa sequence is more
successful than BLAST searches directly against the moa genome to retrieve regions of interest.
File: map_droNov-anoDid_coordinates.pl is a utility script written in Perl programming language to retrieve
moa sequence for genomic regions of interest based on an input set of coordinates in BED format for the
refernce emu genome assembly.
Users should ensure that Perl and Bioperl are installed on their system, and should edit directory
names on lines 39 & 51 if working in a non-UNIX style environment.
To run the script, provide a BED format input file that lists regions to retrieve
(an example input file for testing purposes is provided: example_coordinates.bed).
The input file is in BED format, i.e. tab-delimited text with:
Column 1: scaffold ID
Column 2: emu start position (0-based coordinate)
Column 3: emu end position (1-based coordinate)
Column 4: a unique ID that will be used to name output files
Column 5: score (n/a, put . as placeholder)
Column 6: strand (+ or -, if - is specified, output alignments will be reverse complemented
relative to orientation in the emu reference genome)
The script should be run in a directory that contains the (decompressed) 'droNov-anoDid_alignments'
subdirectory, and the input BED file of emu coordinates is provided as a command-line parameter, e.g.:
perl map_droNov-anoDid_coordinates.pl example_coordinates.bed
For each specified input region, the script will output the aligned emu-moa sequence to the
region_alignments/ output directory, and the mapped moa coordinates to the output BED format
file 'anoDid_coordinates.bed'. (Note that this output should be renamed or moved if users
don't wish to overwrite it upon subsequent use of script).
The script will warn (to screen) if an input region has no corresponding scaffold alignment.
If there are no typos in the provided input file and all alignment .fastas have been
downloaded & decompressed, this means that there were no moa reads mapped to the missing
emu reference scaffold (and hence it did not appear in the moa genome assembly).
The script will also warn if an input region has no moa bases (e.g. only gaps/Ns), but will
still output the emu-moa alignment subregion.
A final note: gap insertion in the emu-moa alignments follows scoring penalties used in HTS
read mapping and, for instance, more divergent sequences will sometimes be mapped as an
insertion in one species followed by an insertion of equal length in the second species.
Users should therefore review output, and consider whether de novo realignment of the
subregion is more suitable for their intended downstream use (e.g. by removing gaps and
realigning with MAFFT or other multiple sequence alignment software).
#########################################################################
DATA-SPECIFIC INFORMATION FOR: emu-moa_alignments_mapDamage_corrected_assembly.tar.gz
Archive directory structure:
emu-moa_alignments_mapDamage_corrected_assembly
├── droNov-anoDid_alignments
├── example_coordinates.bed
└── map_droNov-anoDid_coordinates.pl
Description:
The directory structure and contents of this this archive are identical to that outlined for
emu-moa_alignments_original_assembly.tar.gz (above), except that it uses the mapDamage corrected
version of the moa nuclear genome assembly.
#########################################################################
DATA-SPECIFIC INFORMATION FOR: MHC_alignments.zip
Archive directory structure:
MHC alignments/
├── MHC_classI_ex3_mafft_subset.fa
├── MHCclassIIex2_tblastn_alignment.fa
├── chicken_classII_ex2_aa_NP_001038144.2_blast_qcov50_mafft_clipkit.fa
└── chicken_classII_ex2_aa_NP_001038144.2_blast_qcov50_mafft_clipkit.fa.tre
Description:
Multiple sequence alignments in FASTA format for major histocompatibility complex (MHC) class Ia gene exon 3
(MHC_classI_ex3_mafft_subset.fa) and MHC class IIB gene exon 2 sequences (MHCclassIIex2_tblastn_alignment.fa
and chicken_classII_ex2_aa_NP_001038144.2_blast_qcov50_mafft_clipkit.fa), and phylogenetic tree reconstruction
in Newick format for MHC class IIB exon 2 sequences (chicken_classII_ex2_aa_NP_001038144.2_blast_qcov50_mafft_clipkit.fa.tre)
#########################################################################
DATA-SPECIFIC INFORMATION FOR: microsatellites.tar.gz
Archive directory structure:
microsatellites
├── mapDamage_corrected_assembly
│ ├── dinucleotide_repeats
│ │ ├── dinucleotide_bam
│ │ └── dinucleotide_msa
│ └── trinucleotide_repeats
│ ├── trinucleotide_bam
│ └── trinucleotide_msa
└── original_assembly
├── dinucleotide_repeats
│ ├── dinucleotide_bam
│ └── dinucleotide_msa
└── trinucleotide_repeats
├── trinucleotide_bam
└── trinucleotide_msa
Description:
Top-level subdirectories are included for data sets that use moa sequence from the original (uncorrected)
genome assembly, and from the mapDamage corrected assembly.
Separate subdirectories are included for di- and trinucleotide microsatellite repeats.
Within each subdirectory is a further 'bam' subdirectory containing the indel-realigned reads mapped to each
microsatellite region (repeat + flanking sequence) in BAM format, and the corresponding reference allele
genomic sequence in FASTA format.
A second 'msa' subdirectory contains multiple sequence alignments of the little bush moa
seqence for each microsatellite aligned with the corresponding genomic region in other ratite species, in FASTA format.
#########################################################################
DATA-SPECIFIC INFORMATION FOR: moa_mtDNA_CR_phylogeny.tar.gz
Archive directory structure:
moa_mtDNA_CR_phylogeny/
├── Moa_mtDNA_CR_alignment.fasta
└── Moa_mtDNA_CR_tree.tre
Description:
File: Moa_mtDNA_CR_alignment.fasta
Alignment in FASTA format of mitochondrial control region (CR) sequence with previously published
CR sequences for moa from Bunce et al. (2009, PNAS 106:20646-20651).
File: Moa_mtDNA_CR_tree.tre
Maximum-likelihood phylogenetic tree in Newick format resulting from analysis of Moa_mtDNA_CR_alignment.fasta
#########################################################################
DATA-SPECIFIC INFORMATION FOR: OR_genes.zip
Archive directory structure:
OR_genes
├── all_unique_ORs.txt
├── crocOR14+unique3+chick_OR14_probe_aligned_gt05.phy
├── crocOR14+unique3+chick_OR14_probe_aligned_gt05.phy.treefile
└── groupInfo_list_ORtree_1833_taxa.RDS
Description:
File: all_unique_ORs.txt
Final list of olfactory receptor (OR) genes and genomic coordinates following curation to remove
sequences below a minimum length threshold and to remove duplicated loci.
File: crocOR14+unique3+chick_OR14_probe_aligned_gt05.phy
Multiple sequence alignment in PHYLIP format of the 1833 OR loci identified from TBLASTN similarity searches.
File: crocOR14+unique3+chick_OR14_probe_aligned_gt05.phy.treefile
Maximum-likelihood phylogeny in Newick format resulting from analysis of crocOR14+unique3+chick_OR14_probe_aligned_gt05.phy.
File: groupInfo_list_ORtree_1833_taxa.RDS
R object containing list of OR genes for each taxon in the OR tree in Fig. 3, as well as the list of tip names in the phylogenetic tree of OR genes for each clade of OR indicated in Fig. 3.
########################################################################
DATA-SPECIFIC INFORMATION FOR: parse_mapDamage_pileup.pl
Description: Accessory Perl script used to call the final mapDamage corrected nuclear genome sequence by masking
bases (by Ns) from the original moa assembly that no longer had coverage at min. base quality BASEQ >= 20 following
base quality recalibration using mapDamage2 (Jonsson et al. 2013, Bioinformatics 29: 1682-1684).
########################################################################
DATA-SPECIFIC INFORMATION FOR: pop_gen.zip
Archive directory structure:
pop_gen
├── high_pi_windows.tsv
├── moa.genome
├── moa.windowed.pi
├── moa.windows.covered
├── moa_bychr.tsv
└── process_popgen.R
Description:
File: high_pi_windows.tsv
Tab-delimited text file listing 22 genomic windows of 10 Kb that showed elevated values for nucleotide diversity (pi).
File: moa.genome
Tab-delimited text file of moa scaffold accessions and nucleotide lengths.
File: moa.windowed.pi
Tab-delimited text file listing values of nucleotide diversity (pi) for 10 Kb sized windows across the moa genome assembly.
File: moa.windows.covered
Tab-delimited text file listing coverage of mapped reads across 10 Kb windows of the moa genome assembly.
File: moa_bychr.tsv
Tab-delimited text file listing the number of single nucleotide polymorphisms (SNPs) and theta for each
scaffold in the moa genome assembly, used for the estimation of effective population size (Ne).
File: process_popgen.R
Script in R programming language used for the analysis and graphing of population genetic parameters contained in
the tab-delimited text files outlined above.
########################################################################
DATA-SPECIFIC INFORMATION FOR: Sackton_et_al._2019_update.tar.gz
Archive directory structure:
Sackton_et_al._2019_update
├── moa_mapDamage_coordinates
│ ├── anoDid_mapDamage_CNEE_coordinates.bed.gz
│ └── anoDid_mapDamage_exon_coords.bed.gz
└── paleognath_phylogeny
├── MP-EST_species_trees
│ ├── CNEEs_mp-est_species_tree.tre
│ ├── TENT_mp-est_species_tree.tre
│ ├── UCEs_mp-est_species_tree.tre
│ └── introns_mp-est_species_tree.tre
├── RAxML_gene_trees
│ ├── RAxML_bestML_trees
│ │ ├── CNEEs.tar.gz
│ │ ├── UCEs.tar.gz
│ │ └── introns.tar.gz
│ └── RAxML_bootstraps
│ ├── CNEEs.tar.gz
│ ├── UCEs.tar.gz
│ └── introns.tar.gz
└── alignments
├── CNEEs.tar.gz
├── UCEs.tar.gz
└── introns.tar.gz
Description:
Top-level subdirectory moa_mapDamage_coordinates/ contains two additional files added as an update
to the Sackton et al. 2019 manuscript data release (Science 364: 74-78). The original data release
(available from Dryad Digital Repository DOI: https://doi.org/10.5061/dryad.v72d325) included moa genome
coordinates for conserved nonexonic elements (CNEEs) and protein-coding exons for the
original moa genome assembly only. Here, we provide the genomic coordinates for the
mapDamage corrected assembly in BED format. (Note that sequences in the Sackton et al.
data release are unaffected, it was simply that the mapDamage coordinate BEDs were not
included).
Top-level subdirectory paleognath_phylogeny/ contains data for phylogenetic inference of paleognaths from data sets
incorporating the mapDamage corrected moa sequence. Included are subdirectories containing
multiple sequence alignments in FASTA format, RAxML gene trees (both the 'bestML', or
optimal gene trees as well as trees inferred from 500 bootstrap replicates per locus), and the
MP-EST species trees for each marker type as well as the total evidence nuclear tree (TENT)
combining all loci. Refer to Sackton et al. 2019 (Science 364:74-78) or Cloutier et al. 2019
(Syst. Biol. 68:937-955) for a full description of methods.
########################################################################
DATA-SPECIFIC INFORMATION FOR: sensory_gene_alignments.zip
Archive directory structure:
sensory_gene_alignments
├── Archosaur_alignment_positive_selection_tests_trimmed.fasta
├── Large_T2R_alignment_for tree_SupplementFigure.fasta
└── Opsin_alignment_positive_selection_tests_trimmed.fasta
Description:
Multiple sequence alignments in FASTA format for bitter taste receptors (T2Rs, files
Archosaur_alignment_positive_selection_tests_trimmed.fasta and Large_T2R_alignment_for tree_SupplementFigure.fasta),
and opsin visual pigment genes (file Opsin_alignment_positive_selection_tests_trimmed.fasta).
########################################################################
DATA-SPECIFIC INFORMATION FOR: species_codes.txt
Description: Tab-delimited text file listing the scientific and common names, and the 6-letter code for
species used in the candidate gene, microsatellite, and whole-genome phylogenomics analyses included in
this data release.
########################################################################
DATA-SPECIFIC INFORMATION FOR: TE_analysis.zip
Archive directory structure:
TE analysis
├── dnaPipeTE_results
│ ├── aptrow_reads_per_component_and_re_annotation_deepTE
│ ├── chicken_reads_per_component_and_re_annotation_deepTE
│ ├── emu_reads_per_component_and_re_annotation_deepTE
│ ├── eudromia_reads_per_component_and_re_annotation_deepTE
│ ├── moa_reads_per_component_and_annotation
│ └── ostrich_reads_per_component_and_re_annotation_deepTE
└── moa_unannotated_TEs
└── moa_unannotated_deepTE_blast_to_ratites_chick_GGswu_by_species.txt
Description:
Top-level subdirectory dnaPipeTE_results/ contains output files in tab-delimited text format
using dnaPipeTE to estimate the genomic fraction of transposable elements (TEs) in representative
palaeognath genomes and the outgroup chicken.
Top-level subdirectory moa_unannotated_TEs/ contains a tab-delimited text file
(moa_unannotated_deepTE_blast_to_ratites_chick_GGswu_by_species.txt) that summarizes
the results of BLAST similarity searches of unannotated moa repeats from the dnaPipeTE
analysis queried against other representative paleognath genomes as well as chicken and
Saltwater crocodile outgroups.
Sharing/Access information
1. Licenses/restrictions placed on the data: CC0 1.0 Universal (CC0 1.0) Public Domain
2. Links to publications that cite or use the data:
Edwards, S. V., Cloutier, A., Cockburn, G., Driver, R., Grayson, P., Katoh, K., Baldwin, M. W., Sackton, T. B., Baker, A. J. (2024). A nuclear genome assembly of an extinct flightless bird, the little bush moa. Science Advances: adj6823.
3. Links to other publicly accessible locations of the data:
Aotearoa Genomic Data Repository, Project ID TAONGA-AGDR00049 Project DOI https://doi.org/10.57748/M42Z-SW23 (nuclear and mitochondrial genome assemblies)
National Center for Biotechnology Information (NCBI), BioProject PRJNA433423 GenBank assembly GCA_006937325.1 (original nuclear genome assembly), GenBank MK778441 (mitochondrial genome assembly), Short Read Archive accession SRP132423 (raw sequencing reads); BioProject PRJNA534317 GenBank assembly GCA_006938045.1 (mapDamage corrected nuclear genome assembly)
4. Links/relationships to ancillary data sets: None
5. Was data derived from another source? No
A. If yes, list source(s): NA
6. Recommended citation for this dataset:
Cloutier, A., Edwards, S. V., Cockburn, G., Driver, R., Grayson, P., Katoh, K., Baldwin, M. W., Sackton, T. B., Baker, A. J. (2024). Data from: A nuclear genome assembly of an extinct flightless bird, the little bush moa. Dryad Digital Repository. https://doi.org/10.5061/dryad.d51c59zxp
- Edwards, Scott V.; Cloutier, Alison; Cockburn, Glenn et al. (2024). A nuclear genome assembly of an extinct flightless bird, the little bush moa. Science Advances. https://doi.org/10.1126/sciadv.adj6823
