Data from: Assessing the validity of the fixed tree topology assumption in phylodynamic inference
Abstract
Fixed tree topologies are widely used in phylodynamic analyses to reduce computational burden, yet the consequences of this assumption remain insufficiently understood. Here, we systematically assess the impact of various fixed-topology strategies on phylogenetic and phylodynamic parameter estimates across a diverse set of viral datasets. We compare fully Bayesian joint inference with fixed-topology strategies, including conditioning on maximum likelihood trees subsequently dated with LSD or TreeTime. Our analyses show that global parameters of the substitution and site models are largely robust to the fixed-topology assumption, whereas parameters that depend on the temporal structure of the tree, such as molecular clock rates, node ages, and demographic histories, can exhibit substantial biases. We do treat unconstrained Bayesian analyses as the reference, although we recognize that these too are model-based approximations. Nevertheless, our results highlight serious discordance associated with fixing the topology and underscore the need for faster, time-aware methods that simultaneously integrate topology and parameter estimation. These findings raise important questions about the balance between computational efficiency and inferential accuracy in phylodynamic studies.
Mathieu Fourment, Jiansi Gao, Marc A Suchard, Frederick A Matsen IV. Assessing the Validity of the Fixed Tree Topology Assumption in Phylodynamic Inference. arXiv:2511.22005 (2025)
1. Overview of the deposit
Phylodynamic analyses infer epidemiological and evolutionary quantities — such as the age of the common ancestor of an outbreak, the rate at which the pathogen evolves, and how the effective population size of the pathogen changed through time — from time-stamped genetic sequences.
These quantities are estimated jointly with a phylogenetic tree, which is expensive to compute, so a common shortcut is to fix the tree topology to a single point estimate obtained beforehand.
The associated manuscript quantifies how much this shortcut changes the resulting estimates.
This deposit contains the inputs needed to reproduce those analyses for 15 published viral datasets: the configuration files that define each Bayesian phylogenetic analysis, plus, for the datasets that cannot be redistributed, the accession numbers needed to retrieve the sequences.
It does not contain the MCMC output of the analyses (posterior samples of trees and parameters), which amounts to several gigabytes.
Two files are deposited:
| File | Contents |
|---|---|
data.zip |
The input files, organised in one folder per dataset. |
2. Organisation of data.zip
data.zip unpacks to a single data/ directory containing 15 sub-directories, one per dataset.
Every sub-directory follows the same pattern, so the contents are described once here and in detail in the other sections:
data/
├── <dataset>/
│ ├── run01.xml # BEAST v1.10.5 configuration file: sequences, dates, models, MCMC settings
│ └── gisaid.txt # GISAID accession numbers; present only for the 3 SARS-CoV-2 datasets
Throughout this README, <dataset> stands for any one of the 15 folder names listed in Section 3.
3. The 15 datasets
Folder names follow the pattern <virus/gene>_<first three letters of the first author of the source study><two-digit year of that study>; for example lassaL_kli22 is the Lassa virus L-segment dataset taken from a 2022 study whose first author's surname begins with "Kli". Each dataset was assembled by the authors of the original study; we re-used their sequence alignments, sampling dates and model specifications without modification, so that the models below are those of the source publications rather than choices made for this manuscript.
Folder: ebov_dud17
- Pathogen / genomic region: Ebola virus
- Sequences: 1610
- Alignment length (nt):18992
- Sampling dates (decimal years): 2014.21 – 2016.00
- Data partitions: 4 (codon positions 1, 2, 3 of sites 1–14517 + sites 14518–19028)
- Substitution model: HKY+Γ per partition
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 2.0 y
Folder: ebov_mba21
- Pathogen / genomic region: Ebola virus
- Sequences: 297
- Alignment length (nt): 16757
- Sampling dates (decimal years): 2018.57 – 2020.13
- Data partitions: 2 (codon positions 1+2, 3)
- Substitution model: HKY+Γ per partition
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 1.75 y
Folder: fluH1L_bed15
- Pathogen / genomic region: Influenza A/H1N1
- Sequences: 2144
- Alignment length (nt): 1695
- Sampling dates (decimal years): 2000.00 – 2011.00
- Data partitions: 2 (codon positions 1+2, 3)
- Substitution model: HKY per partition
- Molecular clock: strict
- Coalescent tree prior: constant size
Folder: fluH3L_bed15
- Pathogen / genomic region: Influenza A/H3N2
- Sequences: 4006
- Alignment length (nt): 1698
- Sampling dates (decimal years): 2000.00 – 2013.00
- Data partitions: 2 (codon positions 1+2, 3)
- Substitution model: HKY per partition
- Molecular clock: strict
- Coalescent tree prior: constant size
Folder: fluPb2_wor14
- Pathogen / genomic region: Influenza A, PB2 segment
- Sequences: 354
- Alignment length (nt): 2277–2280
- Sampling dates (decimal years): 1918.00 – 2011.00
- Data partitions: 2 (codon positions 1+2, 3)
- Substitution model: HKY+Γ per partition
- Molecular clock: fixed local clock, 9 host-associated clades
- Coalescent tree prior: Skygrid, 50 intervals, 150 y
Folder: fluVicL_bed15
- Pathogen / genomic region: Influenza B/Victoria
- Sequences: 1999
- Alignment length (nt): 1755
- Sampling dates (decimal years): 2000.00 – 2013.00
- Data partitions: 2 (codon positions 1+2, 3)
- Substitution model: HKY per partition
- Molecular clock: strict
- Coalescent tree prior: constant size
Folder: hiv_far14
- Pathogen / genomic region: HIV-1 group M
- Sequences: 927
- Alignment length (nt): 438
- Sampling dates (decimal years): 1959.00 – 2005.00
- Data partitions: 1
- Substitution model: GTR+Γ
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 90 y
Folder: lassaL_kli22
- Pathogen / genomic region: Lassa virus, L segment
- Sequences: 551
- Alignment length (nt): 7038
- Sampling dates (decimal years): 1969.00 – 2019.18
- Data partitions: 1
- Substitution model: GTR+Γ
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 500 y
Folder: mumps_mon21
- Pathogen / genomic region: Mumps virus
- Sequences: 467
- Alignment length (nt): 15393
- Sampling dates (decimal years): 2006.00 – 2018.09
- Data partitions: 1
- Substitution model: HKY+Γ
- Molecular clock: strict
- Coalescent tree prior: Skygrid, 50 intervals, 25 y
Folder: rabies_via23
- Pathogen / genomic region: Rabies virus
- Sequences: 290
- Alignment length (nt): 10815 (coding) + 1068 (non-coding)
- Sampling dates (decimal years): 1997.02 – 2016.07
- Data partitions: 3 (codon positions 1+2, 3 of the coding alignment; non-coding alignment)
- Substitution model: GTR+Γ per partition
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 20 y
Folder: sars2_can20
- Pathogen / genomic region: SARS-CoV-2
- Sequences: 1046
- Alignment length (nt): not distributable
- Sampling dates (decimal years): 2019.98 – 2020.33
- Data partitions: 1
- Substitution model: HKY+Γ
- Molecular clock: strict
- Coalescent tree prior: Skygrid, 20 intervals, 0.5 y
Folder: sars2_lem21
- Pathogen / genomic region: SARS-CoV-2
- Sequences: 3241
- Alignment length (nt): not distributable
- Sampling dates (decimal years): 2020.08 – 2020.83
- Data partitions: 1
- Substitution model: HKY+Γ
- Molecular clock: strict
- Coalescent tree prior: Skygrid, 20 intervals, 0.8 y
Folder: sars2_pek22
- Pathogen / genomic region: SARS-CoV-2
- Sequences: 717
- Alignment length (nt): not distributable
- Sampling dates (decimal years): 2019.98 – 2020.12
- Data partitions: 1
- Substitution model: GTR+I
- Molecular clock: strict
- Coalescent tree prior: Skygrid, 20 intervals, 0.37 y
Folder: wnv_del20
- Pathogen / genomic region: West Nile virus
- Sequences: 801
- Alignment length (nt): 10302
- Sampling dates (decimal years): 1999.00 – 2016.65
- Data partitions: 1
- Substitution model: GTR+Γ
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 19 y
Folder: zikv_gru19
- Pathogen / genomic region: Zika virus
- Sequences: 283
- Alignment length (nt): 10269
- Sampling dates (decimal years): 2013.81 – 2018.05
- Data partitions: 3 (codon positions 1, 2, 3)
- Substitution model: HKY+Γ per partition
- Molecular clock: uncorrelated relaxed (lognormal)
- Coalescent tree prior: Skygrid, 50 intervals, 5 y
Column notes:
- Sequences — number of aligned sequences (tips of the tree).
- Sampling dates — decimal years (e.g. 2020.5 is mid-2020); larger values are more recent.
- Data partitions — subsets of alignment columns that are given their own substitution model
parameters. "Codon positions 1+2, 3" means first and second positions of each codon share one set
of parameters and third positions (which evolve faster) get another. All partitions always share
the same tree. - Substitution model — HKY distinguishes transitions from transversions; GTR allows all six
reversible nucleotide exchange rates to differ; "+Γ" adds gamma-distributed rate variation across
sites (4 discrete rate categories). - Molecular clock — how substitution rates vary across branches: strict = one rate everywhere;
uncorrelated relaxed (lognormal) = each branch draws its own rate from a shared lognormal
distribution; fixed local clock = separate rates for pre-specified clades (here, influenza
lineages circulating in different hosts). - Coalescent tree prior — the demographic model that supplies the prior on the timed tree.
Constant size assumes an unchanging effective population size; the Skygrid is a
non-parametric model that estimates effective population size in a fixed number of time intervals
going back to a specified cut-off (the last two numbers in the column), smoothed by a Gaussian
Markov random field.
4. run01.xml — BEAST XML configuration files
4.1 Purpose
run01.xml is a configuration file in the XML format of BEAST v1.10.5 (Bayesian Evolutionary
Analysis Sampling Trees). One such file completely specifies a Bayesian phylodynamic analysis: the data, the models, the prior distributions, the MCMC settings, and the output to be written. BEAST is run non-interactively on the file, e.g.:
beast run01.xml
and samples from the joint posterior distribution of the timed phylogenetic tree, the substitution
model parameters, the clock rate(s) and the demographic (effective population size) trajectory.
In this study these files are the starting point for two comparisons. The deposited file is the
free-topology analysis, in which the tree topology is estimated together with everything else. The
fixed-topology counterpart used in the manuscript is generated programmatically from the same file
by the pipeline, which (i) inserts a pre-computed tree as a <newick id="startingTree">
element in place of the random starting-tree simulator, and (ii) deletes the four operators that
propose changes to the topology (subtreeSlide, narrowExchange, wideExchange,
wilsonBalding). Everything else — data, priors, clock and demographic models — is identical, so
any difference between the two posteriors is attributable to fixing the topology.
4.2 Files produced when run01.xml is executed
These outputs are not part of the deposit; they are what a user obtains by running the file.
| Output file | Written every | Contents |
|---|---|---|
prob_run01.log |
100 states | Log posterior, log prior, log likelihood and log coalescent (Skygrid) density. |
param_run01.log |
100 states | Substitution parameters (kappa or gtr.rates, frequencies, alpha), clock parameters (clock.rate or ucld.mean/ucld.stdev, meanRate, coefficientOfVariation, covariance), skygrid.precision, rootHeight and treeLength. |
sky_run01.log |
100 states | The skygrid.logPopSize vector — the log effective population size in each time interval, i.e. the sampled demographic trajectory. |
opsTipheight_run01.log |
100 states | Sampled tip ages; present only for the datasets with uncertain sampling dates. |
run01.trees |
10 000 states | Posterior sample of timed trees in NEXUS format, annotated with branch rates. |
run01.ops |
end of run | Acceptance-rate report for every operator, used to diagnose MCMC performance. |
Tab-delimited .log files are readable with Tracer; .trees files
with TreeAnnotator, FigTree, DendroPy, ape or ETE.
5. gisaid.txt and access to the SARS-CoV-2 sequences
GISAID's terms of use do not permit sequence data to be redistributed, so the three SARS-CoV-2
run01.xml files are deposited with every <sequence> element reduced to a single ? placeholder
character. Everything else in those files — taxon names, sampling dates, models, priors, operators,
MCMC settings — is complete and unmodified.
gisaid.txt is a plain-text, comma-separated list of the GISAID accession numbers (EPI_ISL_######)
of the sequences used. To reconstruct a runnable file, download the corresponding sequences from
gisaid.org (registration required), align them, and paste each aligned
sequence into the <sequence> element of the matching taxon.
| Dataset | Accessions in gisaid.txt |
Taxa in run01.xml |
How taxon names map to accessions |
|---|---|---|---|
sars2_can20 |
742 | 1046 | Taxon names are hCoV-19/…|EPI_ISL_######|date; the accession is the second field. The remaining 304 taxa carry a location instead of an accession (hCoV-19/…|Brazil/SP/SaoPaulo|date) and are sequences generated by the source study, which must be obtained from that study. |
sars2_lem21 |
3241 | 3241 | Taxon names are hCoV-19/…|EPI_ISL_######|country|date; one accession per taxon. |
sars2_pek22 |
717 | 717 | Taxon names are Location_MMDDYY_######_Lineage; the six-digit third field is the numeric part of the accession, e.g. HongKong_020720_416314_B ↔ EPI_ISL_416314. |
6. Sharing and access
All files in data.zip may be reused under the licence attached to this Dryad deposit, with the exception of the SARS-CoV-2 sequence data, which is not included and must be obtained from GISAID under its own terms.
The sequence alignments for the other 12 datasets were published by the source studies listed in the manuscript and are redistributed here within the run01.xml files.
