Host hybridization enabled the emergence of a reassorted hantavirus lineage
Data files
Jul 15, 2026 version files 57.87 MB
-
README.md
5.29 KB
-
TULV_EST.N.sce
7.45 MB
-
TULV-CEC-1.sce
8.93 MB
-
TULV-CEC-2.sce
8.08 MB
-
TULV-CEN.N-1.sce
8.71 MB
-
TULV-CEN.N-2.sce
7.46 MB
-
TULV-EST.S-Ref.sce
7.46 MB
-
TULV-EST.S.sce
9.77 MB
Abstract
The exchange of genetic material between individuals is a key driver of evolution and diversification across most branches of life. Segmented viruses can exchange genetic material through reassortment of genomic segments. New viral strains that emerge from reassortments can have greater infection ranges and higher virulence, although concrete examples of the adaptive advantages of reassortants in nature apart from influenza remain rare. We studied here the evolutionary history and consequences of reassortment in Tula orthohantavirus (TULV) in a hybrid zone between evolutionary lineages of its reservoir host, the common vole (Microtus arvalis). Across 58 trapping sites and 127 infected voles, we detected 27 TULV reassortants in a 12.5 km broad zone at the contact of the parental TULV clades, resembling a viral hybrid zone concordant with the hosts’. Phylogenomic analyses revealed three independent reassortment events, but most of the host hybrid zone was dominated by a single strain with a reassorted M-segment, which encodes the surface glycoprotein. We detected clade-specific variation in the glycoprotein’s N-terminal region consisting of five residues, two of which showed evidence of positive selection. In silico 3D modeling of seven glycoproteins confirmed that this N-terminal region has a unique and specific structure for each TULV clade and the dominant reassortants and is the only structurally variable region of the TULV glycoprotein. Our findings suggest that reassortment between the parental TULV clades in the contact region has resulted in a transgressive virus phenotype potentially adapted to hybrid hosts. This demonstrates the potential of zones of hybridization for the emergence of new virus strains with novel evolutionary trajectories.
Dataset DOI: 10.5061/dryad.jdfn2z3qm
Dataset Summary
This dataset contains structural protein models generated to evaluate how host hybridization influences the emergence and structural variability of reassorted hantavirus lineages. Specifically, these models focus on Tula virus (TULV) variants isolated from distinct host backgrounds. By structural modeling of these viral proteins, this study contextualizes structural changes and molecular compatibility governing genetic exchange and reassortment events within hybridizing host populations.
Description of the data and file structure
This repository contains 7 structural coordinate files saved in the .sce (Scene) format, which are configured for native rendering, visualization, and molecular analysis.
File Naming Convention
The files follow a standardized nomenclature reflecting the specific viral lineages and reassortments:
- TULV: Tula virus lineage.
- EST.S / EST.N / CEN.N / CEC: Specific viral clades and reassortments, accounting for distinct spatial variants (e.g., EST.S representing the Southern variant of the Eastern clade and CEC represents a reassortant between CEN.N (S- and L-segment) and EST.N (M-Segment)).
- Ref: Designates the reference structure built from the first published TULV reference genome (Moravia).
File: TULV-EST.S.sce
Description: Protein model corresponding to sample MarCzKe01
File: TULV-CEN.N-1.sce
Description: Protein model corresponding to sample MarDAd01
File: TULV-CEC-1.sce
Description: Protein model corresponding to sample MarDGc09
File: TULV-CEN.N-2.sce
Description: Protein model corresponding to sample MarDPo02
File: TULV-CEC-2.sce
Description: Protein model corresponding to sample MarDPo03
File: TULV_EST.N.sce
Description: Protein model corresponding to sample MarDTd03
File: TULV-EST.S-Ref.sce
Description: Protein model corresponding to the Tulv reference genome Moravia
Key to Abbreviations & Codes
.sce: YASARA Scene file. It preserves the complete 3D structural environment, including atomic coordinates, secondary structure assignments, residue coloring, and camera viewpoints. These files represent complete, fully resolved three-dimensional structural models; no structural coordinates or sequence elements are omitted or coded as missing.
Sharing/Access information
Links to other publicly accessible locations of the data:
- Associated genomic sequences are deposited alongside the primary publication metadata.
Data was derived from the following sources:
- Viral sequence templates isolated from natural host populations at designated contact and hybridization zones.
Code/Software
Software Requirements
All .sce files included in this submission require YASARA (Yet Another Scientific Artificial Reality Application) software to be properly opened, viewed, animated, and edited.
1. Accessing via YASARA (Recommended)
To view these files exactly as they were rendered by the authors, you can use the free runtime visualization tier of the software:
- Software: YASARA View (Freely available for all operating systems).
- Download Link: http://www.yasara.org/products.htm
- How to Open:
- Download, install, and launch YASARA View.
- Go to
File>Load>YASARA Scene. - Navigate to and select the chosen
.scefile. Alternatively, you can simply drag and drop the.scefile directly into an open YASARA window. - The default view shows all structures rendered as part of the protein modeling. Only "6ZJM-~" models were evaluated for structural comparissons. Individual model variants can be viewed by toggling "Vis" to "No" for all but the desired model.
2. Open-Source Alternatives & Data Interoperability
As all data in Dryad are dedicated to the public domain without restriction, users who prefer not to use proprietary software can access the underlying structural data using open-source packages.
Because the .sce extension is a proprietary scene wrapper, popular open-source molecular viewers cannot natively parse the raw scene files. However, the data can be converted to the standard open-source PDB (Protein Data Bank) format:
- Recommended Open-Source Viewers: PyMOL (Open-Source/Community Edition), UCSF Chimera / ChimeraX, or VMD (Visual Molecular Dynamics).
- How to Extract to PDB Format:
- Option A (Via YASARA View): Open the
.scefile in the free YASARA View application, then selectFile>Save as>PDB fileto write the standard structural coordinates out to a universally readable.pdbfile. - Option B (Alternative Translation): If you do not have YASARA installed, you can utilize open-source molecular file conversion tools or scripts capable of parsing YASARA object elements (e.g., specific workflows leveraging OpenBabel translation configurations) to isolate the standard
ATOMandHETATMcoordinate records from the scene data.
- Option A (Via YASARA View): Open the
Homology modeling, structural alignments and molecular dynamics simulations
To assess structural differences in the M-segment glycoprotein among TULV clades and reassortants, we conducted homology modelling and 3D visualization of the TULV M-segment and its N-terminal ectodomain. We selected six M segment sequences which reflect the full diversity spectrum for AA in the N-terminal ectodomain from the Saxony transect, and the original TULV genome Moravia (Kukkonen et al., 1998) as a TULV-EST.S reference. Homology modelling and molecular dynamics (MD) simulations were performed in YASARA Structure version 20 using adapted macros and structures implemented from (Krieger & Vriend, 2014, 2015). Homology models were built using the standard “hm_build.mcr” macro based on the hantavirus glycoprotein structure “6ZJM” template (Serris et al., 2020). Resulting homology models displayed three different subunit configurations (four in Configuration 1, two in Configuration 2, one in Configuration 3). These different configurations likely reflect stochastic variations in fitting or different states of the protein, as the hantavirus glycoprotein can transition through different structural variations (Rissanen et al., 2017). It is unlikely that these variations reflect actual changes in protein folding, as we even observed different configurations between the near identical TULV-EST.S proteins from the Saxony transect and our reference Moravia (4 AA difference).
Pairwise structural alignments, including RMSDs and sequence identities between the homology models were calculated using MUSTANG (Konagurthu et al., 2006). RMSD of atomic locations across the entire protein tetramer (residue 13-1104) were compared to the RMSD of first nine AA of the Gn-subunits ectodomain (N-terminus). Only the results of the alignment for the four proteins in Configuration 1 are displayed in Supp. Table S10, because structural alignments require identical configuration of all subunits for meaningful evaluation. Individual subunits of the three remaining proteins, when aligned, showed similar levels of low RSMDs, but randomness in the predicted configurations artificially inflates the pairwise RMSDs of the complete proteins. This problem can normally be circumvented by using the configuration of a predicted protein for all future predictions. This was not applicable for our set of proteins, as it would also artificially force the flexible N-terminal ectodomain to change shape to match the reference thus eliminating the key variation between proteins.
MD simulations of all N-terminal regions were performed using the “md.run.mcr” macro and AMBER14 force field, with a simulation duration of 500 ns (Maier et al., 2015). Time simulation steps were set to 1.35 fsec. The simulation box was 'Cube”-shaped and extended at least 10 Å to each side of the model (extension=10), was filled with 0.9% NaCl and the TIP3P water model was used at physiological pH 7.4. Further settings were: temperature at 298K, pressure at 1 bar, density = 0.997, cutoff 8Å- periodic cell boundary and longrange coulomb forces (particle-mesh Ewald). The solute was kept from diffusing and crossing periodic boundaries using the CorrectDrift function (Krieger & Vriend, 2014, 2015). The four flanking amino acids were fixed as anchor point during the MD simulations. To determine structural differences between the regions of interest, root mean square fluctuations (RMSF) were calculated by the “md_analyze” macro. The RMSF is the fluctuation of every heavy atom compared to the mean structure within a simulation cell. The average RMSF of the constituent atoms is used to determine the RMSF per solute residue.
