{"id":774461,"date":"2026-09-18T07:22:32","date_gmt":"2026-09-18T07:22:32","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/774461\/"},"modified":"2026-09-18T07:22:32","modified_gmt":"2026-09-18T07:22:32","slug":"scalable-near-real-time-bayesian-phylogenetics-for-outbreaks-with-delphy","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/774461\/","title":{"rendered":"Scalable near-real-time Bayesian phylogenetics for outbreaks with Delphy"},"content":{"rendered":"<p>For brevity, we summarize the essentials of Delphy\u2019s operation here; full details are provided in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>.<\/p>\n<p>EMATs<\/p>\n<p>Delphy represents trees internally as a collection of explicitly timed nodes and a reference sequence. Each non-root node points to an earlier parent, and each inner node to two later children. Every instant in the tree has an associated sequence, encoded as successive differences from a reference sequence consisting of L states, each in {A,C,G,T}. The root node stores mutations (a,\u2113,b) recording a difference at site \u2113 between the reference state a and the root sequence state b; we refer to these mutations as being above the root node. Every other node is annotated with a sequence of mutations from its parent to itself. Each mutation (a,\u2113,b,t) records a change in state at time t from a to b. Thus, the sequence at point x on the tree is obtained starting with the reference sequence, applying the mutations above the root node to obtain the root sequence, then successively applying in order all mutations on the unique path from the root node to x.<\/p>\n<p>Nodes are also annotated with missations, tuples (a,\u2009\u2113) recording that for all tips downstream of this node, but not its parent, the state of site \u2113 is unknown (missing, \u2018N\u2019 in the input); at the parent, the state is a. The tree topology, mutational history and tip sequences jointly completely determine all missations. Missations are encoded using two complementary structures: an ordered sequence of disjoint, non-consecutive half-open intervals [\u2113start, \u2113end); and a map of sites \u2113 to states a whenever the reference state is not a. This representation reflects that missing data typically appears in a few long gaps, and that the site-to-state map is sparse, as the state at site \u2113 is typically the root state when root-to-tip times are small compared with mutation rates (assuming a reference sequence matching a representative root sequence).<\/p>\n<p>To support uncertain tip dates, each tip has a minimum and maximum time, which coincide when there is no uncertainty.<\/p>\n<p>EMATs have evident consistency requirements for node times, mutation times and identities and missation states (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>).<\/p>\n<p>Posterior distribution<\/p>\n<p>Delphy samples trees \\({\\mathcal{T}}\\) and associated model parameters \u03b8 using MCMC according to the following (unnormalized) posterior distribution:<\/p>\n<p>$$P({\\mathcal{T}},\\theta )\\propto L({\\mathcal{T}}\\,)\\times {\\mathcal{G}}({\\mathcal{T}}|\\theta )\\times {\\pi }_{\\mathrm{anc}}({\\mathcal{T}}|\\theta )\\times {\\pi }_{\\theta }(\\theta )$$<\/p>\n<p>\n                    (1)\n                <\/p>\n<p>The factors are as follows:<\/p>\n<p>\\({\\pi }_{\\theta }(\\theta )\\) is the prior distribution for the model parameters \u03b8.<\/p>\n<p>\\({\\pi }_{{\\rm{a}}{\\rm{n}}{\\rm{c}}}({\\mathcal{T}}\\,|\\theta )\\) is an ancestry prior for the tree\u2019s topology given the model parameters \u03b8.<\/p>\n<p>\\(G({\\mathcal{T}}\\,|\\theta )\\) is a genetic prior for the particular mutational history decorating the EMAT \\({\\mathcal{T}}\\).<\/p>\n<p>\\(L({\\mathcal{T}}\\,)\\) is the likelihood of the data given the EMAT \\({\\mathcal{T}}\\); as Delphy currently restricts tip sequences to be either definite (A,C,G,T) or completely missing (N), this likelihood is simply 1 if the EMAT is consistent (see above) or 0 otherwise.<\/p>\n<p>The first two priors are described below. The genetic prior is the probability that a random root sequence evolved with the evolution model parametrized by \\(\\theta \\) produces the mutational history in EMAT \\({\\mathcal{T}}\\). Explicitly,<\/p>\n<p>$${\\mathcal{G}}({\\mathcal{T}}\\,|\\theta )=\\left\\{\\prod _{{\\ell }\\in \\xi (r)}{\\pi }_{{s}_{r}^{({\\ell })}}^{({\\ell })}\\right\\}\\times \\exp \\,\\left[-{\\int }_{x\\in {\\mathcal{T}}}\\lambda (x){\\rm{d}}x\\right]\\times \\prod _{(a,{\\ell },b)\\in {\\mathcal{M}}}{Q}_{ab}^{({\\ell })}$$<\/p>\n<p>\n                    (2)\n                <\/p>\n<p>where:<\/p>\n<p>\\({Q}_{{ab}}^{({\\ell })}\\) is the rate at which state a at site \u2113 transitions to state b, as parametrized by the evolutionary model parameters in \u03b8.<\/p>\n<p>\\({Q}_{a}^{({\\ell })}=-{Q}_{{aa}}^{({\\ell })}\\) is the rate at which state a at site \u2113 transitions to any other state.<\/p>\n<p>\\({s}_{x}^{({\\ell })}\\) is the state of site \u2113 at point x on \\({\\mathcal{T}}\\).<\/p>\n<p>\\(\\xi (x)\\) is the set of sites for which at least one tip below point x on \\({\\mathcal{T}}\\) is informative (that is, its state is not missing).<\/p>\n<p>\\(\\lambda (x)=\\sum _{{\\ell }\\in \\xi (x)}{Q}_{{s}_{x}^{({\\ell })}}^{({\\ell })}\\) is the sequence-dependent genome-wide mutation rate at point x on \\({\\mathcal{T}}\\) (see the N-pruning discussion below).<\/p>\n<p>\\({\\int }_{x\\in {\\mathcal{T}}}f(x){\\rm{d}}x\\) is an integral of a function \\(f(x)\\) defined at every point on the tree, defined as the sum of time integrals along each branch of \\({\\mathcal{T}}\\).<\/p>\n<p>\\({\\pi }_{a}^{({\\ell })}\\) is the probability that the root sequence has state a at site \u2113.<\/p>\n<p>\\(r\\) is the root node of the tree.<\/p>\n<p>\\({\\mathcal{M}}\\) is the set of all mutations across the tree.<\/p>\n<p>The <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a> further discusses\u00a0the motivation for these choices, considerations for efficient calculations and its equivalence to the standard tree likelihood based on Felsenstein pruning.<\/p>\n<p>Parameter and ancestry priors<\/p>\n<p>Delphy\u2019s priors, \\({\\pi }_{\\theta }(\\theta )\\) and \\({\\pi }_{{\\rm{a}}{\\rm{n}}{\\rm{c}}}({\\mathcal{T}}\\,|\\theta )\\), are currently as follows:<\/p>\n<p>Tip times are, a priori, uniformly distributed between their minimum and maximum times.<\/p>\n<p>The transition rate matrices have the form \\({Q}_{{ab}}^{({\\ell })}=\\mu {\\nu }^{({\\ell })}{q}_{{ab}}\\), where \u03bc is an overall per-site mutation rate, the quantities {\u03bd(\u2113)} are site-relative rates, and \\({q}_{{ab}}\\) are the normalized transition rate matrix elements of an HKY evolution model with transition-transversion rate \u03ba and stationary state frequencies \\({\\pi }_{a}\\). This form implies a strict molecular clock.<\/p>\n<p>By default, the mutation rate \u03bc has an improper uniform prior; a more general gamma prior can also be applied.<\/p>\n<p>The site-relative rates {\u03bd(\u2113)} are either fixed to 1 (no site-rate heterogeneity), or have a priori \\({\\nu }^{({\\ell })} \\sim \\mathrm{Gamma}(\\alpha ,\\alpha )\\), with \\(\\alpha \\sim \\mathrm{Expo}(1)\\) a priori, where ~ means \u2018distributed as\u2019.<\/p>\n<p>The HKY parameters are chosen such that, a priori, \\(\\{{\\pi }_{a}\\} \\sim \\mathrm{Dir}(\\mathrm{1,1,1,1})\\) and \\(\\log (\\kappa ) \\sim N(1,{1.25}^{2})\\).<\/p>\n<p>The ancestry prior is one of:<\/p>\n<p>                  (1)<\/p>\n<p>A standard coalescent prior with an exponentially growing population curve \\(N(t)={n}_{0}{{\\rm{e}}}^{g(t-{t}_{0})}\\), where t0 is the time of the latest tip. By default, the final effective population size n0 has an improper 1\/x prior, while the growth rate g has a Laplace prior with mean 0.001 per year and scale 30.701135 per year. More general priors for n0 and g are also available (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>).<\/p>\n<p>                  (2)<\/p>\n<p>A standard Skygrid flexible population prior<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Gill, M. S. et al. Improving Bayesian population dynamics inference: a coalescent-based model for multiple loci. Mol. Biol. Evol. 30, 713&#x2013;724 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR51\" id=\"ref-link-section-d99178142e4519\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>, where the log of the population curve is specified parametrically at a fixed number of equally spaced times, and at intermediate times is either piecewise constant but discontinuous (staircase, standard) or piecewise linear and continuous (log-linear, particular to Delphy). A priori, the parametrized log-populations follow a random walk with a diffusion constant that allows the population to change by somewhere between halving and doubling in 1\u2009month. We deviate from the original Skygrid by defaulting to fixing instead of inferring this diffusion constant (equivalently, the precision parameter \u03c4), which we find to be both more generally stable and to better encode our intuition of what constitutes reasonable population fluctuations. We also add a penalty for the population curve to assume unrealistically low values (coalescence times below 1\u2009day), which further stabilizes Skygrid for general use. An optional Inverse-Gamma prior can also be applied to the mean population level. Full details are provided in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>.<\/p>\n<p>These choices are suitable for most viral outbreaks, and are currently fixed in Delphy, but their details are not essential. We expect to evolve Delphy to make prior specification more flexible in the future. Most of the above details are the same as the defaults provided by BEAUTi2, and coincide with those used previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Lemieux, J. E. et al. Phylogenetic analysis of SARS-CoV-2 in Boston highlights the impact of superspreading events. Science 371, eabe3261 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR22\" id=\"ref-link-section-d99178142e4535\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>. The priors are discussed further in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>.<\/p>\n<p>MCMC moves<\/p>\n<p>Delphy samples trees and model parameters from the above posterior distribution using MCMC. We distinguish between local moves that affect only a few nodes, and global moves that affect the whole tree.<\/p>\n<p>The following global moves are used (details are provided in the \u2018Global moves\u2019 section of the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>): Gibbs sampling of mutation rate; delta-exchange moves for stationary frequencies \u03c0a and scale moves for \u03ba; scale moves for \u03b1 after integrating out {\u03bd(\u2113)}, followed by Gibbs sampling of {\u03bd(\u2113)}; population parameter moves (exponential model: scale moves on n0 and random walk on g; Skygrid: Gibbs move on overall population size, Hamiltonian Monte Carlo move on log-population sizes and optional Gibbs sampling of \u03c4). The relative simplicity of global moves in the explicit-mutation representation, including the marginalization of {\u03bd(\u2113)} for making moves for \u03b1, was first highlighted in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Lartillot, N. Conjugate Gibbs sampling for Bayesian phylogenetic models. J. Comput. Biol. 13, 1701&#x2013;1722 (2006).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR37\" id=\"ref-link-section-d99178142e4608\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a>.<\/p>\n<p>The following local moves are used (details are provided in the \u2018Local moves\u2019 section of the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>): inner node displacement, branch reform, SPR. For SPR, the regrafting point P\u2032 is proposed using an annealed, approximate mdSPR, scanning points up to 1 mutation away from P in 99% of cases, the whole tree (partition) in 1% of cases. The mutational history on the P\u2032\u2013X branch is a Jukes\u2013Cantor history compatible with end-point sequences, implemented to scale with the number of sequence differences, not the genome size. The interaction of SPR moves with N-pruning is described in full in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>.<\/p>\n<p>Missing data: N-pruning and missations<\/p>\n<p>Missing data substantially complicates using an explicit-mutation representation. Imputing missing data and inferring full mutational histories can be costly in practice (Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>). N-pruning performs Felsenstein pruning only below missations, where the tree likelihood is manifestly 1: the state a below a missation (a,\u2113) evolves to something downstream. This limited Felsenstein pruning preserves the properties of the posterior functional form that Delphy exploits: locality and factorizability; on net, it merely restricts \\({\\ell }\\in \\xi (x)\\) in the above formulas.<\/p>\n<p>Despite the substantial bookkeeping complications (Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>), one can efficiently update missations on trees as subtrees are pruned and regrafted and mutational histories are changed. Missations enter the proposal and acceptance probabilities of all MCMC moves.<\/p>\n<p>Parallelizable coalescent<\/p>\n<p>Delphy augments the usual Kingman coalescent ancestry prior to allow for parallelization (full details are provided in Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>). In particular, terms involving k(t), the number of active lineages at time t, which a priori require a global view of the tree not available within a single partition, are replaced by terms involving kp(t), the number of active lineages in partition p only, and an auxiliary Gaussian coupling field with carefully chosen distribution whose net effect is to recover the Kingman coalescent. For local moves, the Gaussian field is kept fixed, so different kp(t)\u2019s can evolve independently. For global moves, when k(t) is known but static, the coupling field can be Gibbs sampled conditioned on k(t).<\/p>\n<p>To implement this scheme numerically, the Kingman coalescent integral must be discretized at a user-tuneable resolution, below which k(t) and N(t) are approximated as constant. Delphy aims for a discretization with around 400 cells, occasionally changing resolution if the tree height spans too many or too few cells. This default value is often suitable, but can lead to artifacts when there are too many active branches (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). Note also that there is a subtle interaction between the tree partitioning and correct sampling that must be mitigated in concrete implementations (Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">9.1<\/a>).<\/p>\n<p>While the above scheme is correct and permits parallelization, we expect and encourage better parallelization schemes to be developed.<\/p>\n<p>Lineage and mutation prevalence curves<\/p>\n<p>Delphy\u2019s interface shows prevalence curves u(t) for lineages and mutations, equal to the probability that a random member of the population at time t is descended from the subtree below a lineage\u2019s founding inner node, or one below a specific mutation. The coalescent model yields a simple differential equation for u(t), which Delphy solves numerically (full details are provided in Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>). At densely sampled times, where there are many active lineages k(t), then u(t) reduces to the fraction of active lineages with a certain property. More generally, u(t) is an exponentially moving average of that fraction, with decay rate k(t)\/N(t).<\/p>\n<p>Prevalence curves are calculated for each posterior tree, and the mean and 95% HPD range of each such family of curves is displayed to the user.<\/p>\n<p>Automatic detection of burn-in cut-off<\/p>\n<p>Delphy implements a simple heuristic for suggesting an MCMC burn-in cut-off. For a given observable, it calculates the mean and s.d. over the second half of the run, finds the latest time that the observable\u2019s fluctuations exceed 5\u2009s.d., then finds the earliest subsequent time that fluctuations fall within 2\u2009s.d. This time is the suggested cut-off for that observable; Delphy takes the maximum suggested cut-off across the\u00a0log-posterior, mutation rate and total evolutionary time\u00a0traces. Unless a user overrides this suggestion, the cut-off is continuously updated; typically, it stabilizes once the run is well into production.<\/p>\n<p>This heuristic identifies a point near the end of the initial burn-in, then advances to the earliest subsequent \u2018normal\u2019 part of the trace. Empirically, this procedure makes similar choices as a human would. Importantly for Delphy\u2019s accessibility goal, the cut-off suggestion is made automatically: experience with early users showed that requiring manual cut-off selection led to either needless friction or no cut-off at all, biasing the results.<\/p>\n<p>Delphy input formats<\/p>\n<p>Delphy reads multiple-sequence alignments (MSAs) in FASTA or MAPLE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR31\" id=\"ref-link-section-d99178142e4861\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a> format. From each description line, a full sequence ID is extracted after the initial \u2018&gt;\u2019 up to the end of the line. The full ID consists of fields separated by vertical bars (\u2018|\u2019): the first field serves as a short ID for the interface and metadata annotation, the last field is a date specification, and other fields are ignored. A date specification can be an exact date (\u20182025-01-24\u2019), a month (\u20182025-01\u2019), a year (\u20182025\u2019) or a date range (\u20182025-01-20\/2025-01-24\u2019).<\/p>\n<p>Metadata should be a comma-separated value (.csv) or tab-separated value (.tsv) file with a header row. One column should be called \u201cid\u201d or \u201caccession\u201d (case insensitive), with values matching short sequence IDs from the MSA. The remaining columns may have any names and values. Values may be quoted with double-quotes (\u201c), and missing values may be indicated by an empty entry or the values \u2018-\u2019, \u2018noknown\u2019 or \u2018none\u2019 (case insensitive).<\/p>\n<p>Sample MSA and metadata files can be downloaded for the demos on Delphy\u2019s landing page.<\/p>\n<p>APOBEC3-aware evolution model for mpox<\/p>\n<p>Inspired by previous studies<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"O&#x2019;Toole, &#xC1;. et al. APOBEC3 deaminase editing in mpox virus as evidence for sustained human transmission since at least 2016. Science 382, 595&#x2013;600 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR3\" id=\"ref-link-section-d99178142e4883\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Parker, E. et al. Genomics reveals zoonotic and sustained human mpox spread in West Africa. Nature 643, 1343&#x2013;1351 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR4\" id=\"ref-link-section-d99178142e4886\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>, Delphy includes a specialized evolution model suitable for mpox sequences, which can be activated in the web interface under \u201cAdvanced Options\u201d. In brief, each site is classified as having or lacking APOBEC3 context: a site with state C or T has APOBEC3 context when preceded by a T, whereas a site with state G or A has APOBEC3 context when followed by an A. We then use a Jukes\u2013Cantor model with rate \u03bc, modified in sites with APOBEC3 context so C-to-T and G-to-A mutations occur at a rate \u03bc\u2009+\u2009\u03bc*. This setup retains the essence of the models of refs. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"O&#x2019;Toole, &#xC1;. et al. APOBEC3 deaminase editing in mpox virus as evidence for sustained human transmission since at least 2016. Science 382, 595&#x2013;600 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR3\" id=\"ref-link-section-d99178142e4899\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Parker, E. et al. Genomics reveals zoonotic and sustained human mpox spread in West Africa. Nature 643, 1343&#x2013;1351 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR4\" id=\"ref-link-section-d99178142e4902\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> but differs minimally in its details. Full details are provided in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>.<\/p>\n<p>Benchmarks<\/p>\n<p>All benchmarks are in the GitHub and Zenodo data repositories (Data availability). Each benchmark is organized as a series of numbered scripts. Unless noted, the scripts are self-contained and download external data as needed. The repos also include intermediate and final results files, as well as many of the plots here and in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a> (large files are only at Zenodo; and, for GISAID data subject to its data use agreement, we include download instructions but not the actual data). We intend this repository to be executable documentation of every benchmark detail: while all scripts ran correctly at publication, we do not intend to modify them to ensure they continue to run indefinitely.<\/p>\n<p>We used Delphy v.1.1.4 (build 2044, commit a50e378), MAFFT v.7.505 (10 April 2022), BEAST2 v.2.7.7, BEAST X v.10.5.0, Sapling v.0.1.1 (build 2, commit a0b9da1), BEAGLE commit 6480ad3 (Monday, 15 September 2025), IQ-TREE v.2.3.6 and TreeTime v.0.11.4. The sars-cov-2-gisaid-week-by-week and sims benchmarks used Delphy v.1.0 (build 2036, commit 06a7ee4), which lacks a Skygrid population model but is otherwise not materially different. Unless noted, all calculations were run on AWS c7a.2xlarge instances (eight vCPUs), not the web interface (which is 2\u20133\u00d7 slower due to WebAssembly). All Delphy runs were performed twice independently, with convergence checked visually using Tracer<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 56\" title=\"Rambaut, A., Drummond, A. J., Xie, D., Baele, G. &amp; Suchard, M. A. Posterior summarization in Bayesian phylogenetics using Tracer 1.7. Syst. Biol. 67, 901&#x2013;904 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR56\" id=\"ref-link-section-d99178142e4924\" rel=\"nofollow noopener\" target=\"_blank\">56<\/a>.<\/p>\n<p>MCCs were calculated using clade fingerprinting (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>), implemented in the \u2018delphy-mcc\u2019 utility program that is part of Delphy. We verified these match TreeAnnotator2\u2019s output<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Bouckaert, R. et al. BEAST 2.5: an advanced software platform for Bayesian evolutionary analysis. PLoS Comput. Biol. 15, e1006650 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR15\" id=\"ref-link-section-d99178142e4934\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>, which hit severe runtime and memory limits on larger benchmarks; by contrast, delphy-mcc processed our 100,000-sequence simulations in minutes. Delphy-mcc applies a 30% burn-in and behaves as if the option \u2018&#8211;heights ca\u2019 had been given to TreeAnnotator2, so inner node times are the mean tMRCA of the downstream tips over all posterior trees, not just those trees where these tips form a monophyletic clade<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 57\" title=\"Heled, J. &amp; Bouckaert, R. R. Looking for trees in the forest: summary tree from posterior samples. BMC Evol. Biol. 13, 221 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR57\" id=\"ref-link-section-d99178142e4938\" rel=\"nofollow noopener\" target=\"_blank\">57<\/a> (matching the MCC in figure 3a of ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Lemieux, J. E. et al. Phylogenetic analysis of SARS-CoV-2 in Boston highlights the impact of superspreading events. Science 371, eabe3261 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR22\" id=\"ref-link-section-d99178142e4942\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>).<\/p>\n<p>Sampling speed was assessed by dividing ESSs by the wallclock time. For numerical observables, we used LogAnalyser2\u00a0(ref.\u00a0<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Bouckaert, R. et al. BEAST 2.5: an advanced software platform for Bayesian evolutionary analysis. PLoS Comput. Biol. 15, e1006650 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR15\" id=\"ref-link-section-d99178142e4949\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>) with 30% burn-in (verified to suffice by visual inspection of traces, and higher than necessary to avoid subtleties relating to incomplete filtering of burn-in). For tree topology, we implemented the frechetCorrelationESS measure described previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Magee, A., Karcher, M., Matsen, F. A. IV &amp; Minin, V. M. How trustworthy is your tree? Bayesian phylogenetic effective sample size through the lens of Monte Carlo error. Bayesian Anal. 19, 565&#x2013;593 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR58\" id=\"ref-link-section-d99178142e4953\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a>, which quantifies the rate at which pairwise Robinson\u2013Foulds distances between trees tend to their long-term expected value with increasing separation in the run (using clade fingerprinting, we can calculate frechetCorrelationESS for even the large H5N1 benchmark in seconds; Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">11.2<\/a>). In all cases, ESS values were in the hundreds to thousands; exact values are in the benchmark repository.<\/p>\n<p>Clade correlation graphs between any two runs (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2e<\/a>) were computed using a custom program (compare_clades in the data repo), inspired by the analogous CladeSetComparator tool<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Mendes, F. K., Bouckaert, R., Carvalho, L. M. &amp; Drummond, A. J. How to validate a Bayesian evolutionary model. Syst. Biol. 74, 158&#x2013;175 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR47\" id=\"ref-link-section-d99178142e4973\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a> in BEAST 2 but using clade fingerprints (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Information<\/a>). In brief, we first identify all clades appearing in any posterior sample of either run, and filter out any clade with posterior support below 1% in both runs. We then plot for each clade either the posterior support (Support) or the mean clade tMRCA (tMRCAs) in one run versus the other. Error bars show standard errors for posterior support and mean tMRCAs. Dot areas are proportional to the clade size.<\/p>\n<p>SARS-CoV-2 data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Lemieux, J. E. et al. Phylogenetic analysis of SARS-CoV-2 in Boston highlights the impact of superspreading events. Science 371, eabe3261 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR22\" id=\"ref-link-section-d99178142e4985\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a> (sars-cov-2-lemieux)<\/p>\n<p>We obtained accession IDs for the 772 samples in figure 3a of ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Lemieux, J. E. et al. Phylogenetic analysis of SARS-CoV-2 in Boston highlights the impact of superspreading events. Science 371, eabe3261 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR22\" id=\"ref-link-section-d99178142e4992\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a> from the authors (sample_ids.csv), then downloaded the 757 publicly available sequences in GenBank. These were aligned to reference <a href=\"https:\/\/www.ncbi.nlm.nih.gov\/nuccore\/NC_045512.2\" rel=\"nofollow noopener\" target=\"_blank\">NC_045512.2<\/a> (dated to December 2019) using mafft<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 59\" title=\"Katoh, K. &amp; Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol. Biol. Evol. 30, 772&#x2013;780 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR59\" id=\"ref-link-section-d99178142e5003\" rel=\"nofollow noopener\" target=\"_blank\">59<\/a> (&#8211;auto\u2013keeplength), then masked the initial 267 and final 230 sites as in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Lemieux, J. E. et al. Phylogenetic analysis of SARS-CoV-2 in Boston highlights the impact of superspreading events. Science 371, eabe3261 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR22\" id=\"ref-link-section-d99178142e5007\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>.<\/p>\n<p>Delphy was run twice for 2\u2009billion steps (trace every 100,000 steps, posterior trees every 1,000,000). BEAST X and BEAST 2 were run using Delphy\u2019s \u2018equivalent\u2019 XML output without changes for 200\u2009million steps (a rough heuristic: 10 Delphy steps achieve the work of 1 BEAST step). Separate runs were prepared with and without site-rate heterogeneity enabled. Two additional BEAST X runs used K\u2009=\u20092 and K\u2009=\u20098 discrete gamma categories for site-rate heterogeneity.<\/p>\n<p>An unrooted ML tree was built with IQ-Tree 2 (-m HKY+FO or -m HKY+FO+G4), then rooted and dated with TreeTime, which also estimates population growth rates (&#8211;coalescent skyline &#8211;n-skyline 2 &#8211;stochastic-resolve).<\/p>\n<p>MCCs were plotted using baltic library, with inner node metadata inferred through parsimony (ties resolved arbitrarily).<\/p>\n<p>Zika data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Metsky, H. C. et al. Zika virus evolution and spread in the Americas. Nature 546, 411&#x2013;415 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR2\" id=\"ref-link-section-d99178142e5032\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> (zika-metsky-2017)<\/p>\n<p>We extracted the sequences from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Metsky, H. C. et al. Zika virus evolution and spread in the Americas. Nature 546, 411&#x2013;415 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR2\" id=\"ref-link-section-d99178142e5039\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> from the BEAST XML file in its supplementary data (SupplementaryData\/BEAST input and output\/Phylogenetic analyses and model selection\/SRD06-strict-exponential.xml). These were aligned to reference KX197192.1, with the initial 107 and final 428 sites trimmed. Runs and analysis otherwise follow the SARS-CoV-2 benchmark above.<\/p>\n<p>Ebola data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Gire, S. K. et al. Genomic surveillance elucidates Ebola virus origin and transmission during the 2014 outbreak. Science 345, 1369&#x2013;1372 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR1\" id=\"ref-link-section-d99178142e5049\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> (ebola-gire-2014)<\/p>\n<p>We extracted the sequences from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Gire, S. K. et al. Genomic surveillance elucidates Ebola virus origin and transmission during the 2014 outbreak. Science 345, 1369&#x2013;1372 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR1\" id=\"ref-link-section-d99178142e5056\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> from the BEAST XML file in supplementary file\u00a03 (beast\/2014_GN.SL_SRD.HKY_strict_ctmc.exp.xml). In this XML file, the raw sequences are partitioned into genic and intergenic regions, scrambling the mapping to reference KJ660346; we manually deduced the inverse mapping to reconstitute an MSA against this reference. Runs and analysis otherwise follow the SARS-CoV-2 benchmark above.<\/p>\n<p>Ebola data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Dudas, G. et al. Virus genomes reveal factors that spread and sustained the Ebola epidemic. Nature 544, 309&#x2013;315 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR21\" id=\"ref-link-section-d99178142e5066\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a> (ebola-dudas-2017)<\/p>\n<p>We extracted aligned sequences and metadata from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Dudas, G. et al. Virus genomes reveal factors that spread and sustained the Ebola epidemic. Nature 544, 309&#x2013;315 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR21\" id=\"ref-link-section-d99178142e5073\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a> from the companion GitHub repository (Data\/Makona_1610_genomes_2016-06-23.fasta and Data\/Makona_1610_metadata_2016-06-23.csv at <a href=\"https:\/\/github.com\/ebov\/space-time.git\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/ebov\/space-time.git<\/a>, commit 9db59a4). Runs and analysis follow the SARS-CoV-2 benchmark above, except with longer runs (Delphy: 10\u2009billion steps; BEAST X: 1\u2009million) and a Skygrid population model having 24 month-long intervals over the 2\u2009years ending at the latest tip (24 October 2015), with a log-space random walk prior having a 6-month halving\/doubling time (precision \u03c4\u2009=\u200912.3 when effective population sizes are in years).<\/p>\n<p>Mpox data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"O&#x2019;Toole, &#xC1;. et al. APOBEC3 deaminase editing in mpox virus as evidence for sustained human transmission since at least 2016. Science 382, 595&#x2013;600 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR3\" id=\"ref-link-section-d99178142e5093\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> (mpox-otoole-2023)<\/p>\n<p>We extracted the sequences from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"O&#x2019;Toole, &#xC1;. et al. APOBEC3 deaminase editing in mpox virus as evidence for sustained human transmission since at least 2016. Science 382, 595&#x2013;600 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR3\" id=\"ref-link-section-d99178142e5100\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> from a BEAST XML file in its companion GitHub repository (data\/apobec3_2partition.epoch.xml at <a href=\"https:\/\/github.com\/hmpxv\/apobec3\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/hmpxv\/apobec3<\/a>, commit c0b4c9b). We removed one non-public GISAID sequence (EPI_ISL_13983888), and two pre-spillover sequences from before 2017 (KJ642617 from 1971 and KJ642615 from 1978), leaving 41 sequences forming the \u2018hMPXV-1\u2019 ingroup in figure 3c of ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"O&#x2019;Toole, &#xC1;. et al. APOBEC3 deaminase editing in mpox virus as evidence for sustained human transmission since at least 2016. Science 382, 595&#x2013;600 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR3\" id=\"ref-link-section-d99178142e5111\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>. The XML partitions sites into APOBEC3-context and remaining sites; since every site is marked \u2018N\u2019 in at least one partition, the original aligned sequences are trivially reconstituted.<\/p>\n<p>Delphy was run twice for 200\u2009million steps using &#8211;v0-mpox-hack, with trace output every 20,000 steps and posterior trees every 200,000 steps. MCCs and ESSs were calculated as described above.<\/p>\n<p>Mpox data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Parker, E. et al. Genomics reveals zoonotic and sustained human mpox spread in West Africa. Nature 643, 1343&#x2013;1351 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR4\" id=\"ref-link-section-d99178142e5125\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> (mpox-parker-2025)<\/p>\n<p>We extracted the sequences from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Parker, E. et al. Genomics reveals zoonotic and sustained human mpox spread in West Africa. Nature 643, 1343&#x2013;1351 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR4\" id=\"ref-link-section-d99178142e5132\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> from a BEAST XML file in its companion GitHub repository (BEAST\/Mpox_2epoch_combinedDTA.xml.zip at <a href=\"https:\/\/github.com\/andersen-lab\/Mpox_West_Africa\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/andersen-lab\/Mpox_West_Africa<\/a>, commit 2b481da). We removed three non-public GISAID sequences (EPIISL-13953610, EPIISL-13983888 and EPIISL-15008577), and all pre-spillover\/pre-2017 sequences, leaving 177 sequences forming the hMPXV-1 clade in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Parker, E. et al. Genomics reveals zoonotic and sustained human mpox spread in West Africa. Nature 643, 1343&#x2013;1351 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR4\" id=\"ref-link-section-d99178142e5143\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>. As above, the XML partitions sites into APOBEC3-context and remaining, so original sequences are trivially reconstituted.<\/p>\n<p>Delphy was run twice for 1\u2009billion steps using &#8211;v0-mpox-hack, with trace output every 100,000 steps and posterior trees every 1,000,000 steps. MCCs and ESSs were calculated as described above.<\/p>\n<p>For BEAST comparisons, we modified the original BEAST X XML (BEAST\/Mpox_2epoch_combinedDTA.xml.zip) to match Delphy\u2019s sequences and removed phylogeography, spillover detection and pre\/post-spillover partitioning. We shortened the chain to 50\u2009million steps, sufficient for convergence (posterior ESS\u2009=\u2009374). The script to modify this input file and the resulting XML are in our data repository (mpox-parker-2025-beast.xml).<\/p>\n<p>SARS-CoV-2 data from GISAID submitted\/collected by each CDC week (sars-cov-2-gisaid-week-by-week)<\/p>\n<p>We downloaded all SARS-CoV-2 sequences from GISAID collected on or before 31 March 2020 with metadata, filtering out non-human hosts, sequences shorter than 20,000 bases, or with uncertain dates. Each Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> CDC Epiweek panel includes sequences with the submission date in or before that week (filtered to those submitted in 1 December to 28 March 2020, which is the end of CDC Epiweek 2020-13, and collected from 1 December 2019). For Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>, we used collection date instead of submission date (and filtered sequences to those submitted in 1 December 2019 to 31 December 2024, and collected from 1 December 2019). Matching sequences were aligned to <a href=\"https:\/\/www.ncbi.nlm.nih.gov\/nuccore\/NC_045512.2\" rel=\"nofollow noopener\" target=\"_blank\">NC_045512.2<\/a> using mafft, then masked at the initial 268 and final 230 sites (as described previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Lemieux, J. E. et al. Phylogenetic analysis of SARS-CoV-2 in Boston highlights the impact of superspreading events. Science 371, eabe3261 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#ref-CR22\" id=\"ref-link-section-d99178142e5174\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>). No further masking was applied.<\/p>\n<p>Quick initial diagnostic runs revealed clear outlier sequences (for example, inducing a tMRCA to early 2019, lying in an isolated long branch with tens to hundreds of mutations, collection dates far preceding first reported cases in a particular region). After several rounds of iterative refinement by outlier removal, no trees contained obvious outliers. We verified that most offending sequences had been identified as outliers near the beginning of the pandemic, marked as under investigation in GISAID, or appeared in NextStrain or sarscov2phylo exclusion lists. See Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> for the 94 excluded sequences.<\/p>\n<p>Delphy was run on AWS c7a.4xlarge instances (16 vCPUs) for each Epiweek as in the SARS-CoV-2 benchmark, with the following differences: 5\u2009million steps per tip, 10,000 trace and 200 tree samples and at most 1 thread per 100 tips (maximum 32 threads).<\/p>\n<p>Simulated trees for scaling assessment (sims)<\/p>\n<p>We prepared two groups of four SARS-CoV-2-like simulated datasets with N\u2009=\u2009100, 1,000, 10,000 and 100,000 tips. The first group used an exponentially growing population curve, \\(N(t)={n}_{0}\\,{e}^{g(t-{t}_{0})}\\) with n0\u2009=\u20096 years, g\u2009=\u200910 per year and t0\u2009=\u200931 July 2024. The second used a constant population N(t)\u2009=\u2009n0\u2009=\u20092 years. Sample times were drawn from [1 January 2024, 31 July 2024] proportional to N(t), mimicking a uniform sampling of members of the historical viral population, then linked through a standard coalescent simulation. A random 30,000-site root sequence was sampled with \u03c0\u2009=\u2009[0.30,\u20090.18,\u20090.20,\u20090.32] and evolved along the ancestry using the Gillespie algorithm under a site-homogeneous HKY model (mutation rate, 1\u2009\u00d7\u200910\u22123 per site per year, \u03ba\u2009=\u20095). Trees were recorded in Newick format, summaries in JSON, and dated tip sequences as FASTA (N\u2009\u2264\u20091,000) and MAPLE files. To perform these simulations efficiently, we wrote a tool called sapling (Code availability, Data availability and Supplementary Information <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">15.2<\/a>).<\/p>\n<p>Delphy was run twice on a 96-vCPU c7a.24xlarge AWS instance with 5\u2009million steps per tip, 10,000\u2009log and 200 tree samples and at most 1 thread per 20 tips (maximum 192 threads). The n\u2009=\u2009100,000 run also used around 10,000 cells for discretizing the coalescent prior instead of the default about 400 (comparisons at 625\u20135,000 cells are shown in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>).<\/p>\n<p>H5N1 in cattle dataset (h5n1-andersen-2025)<\/p>\n<p>We cloned the Andersen laboratories avian-influenza repository from GitHub (<a href=\"https:\/\/github.com\/andersen-lab\/avian-influenza\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/andersen-lab\/avian-influenza<\/a>) at commit e756a15 (3 October 2025). From 11,541 unique SRRs, we filtered to the 4,080 non-retracted SRRs with a cattle host and B3.13 genotype. For each of H5N1\u2019s 8 segments, we aligned to reference A\/cattle\/Texas\/24-008749-003\/2024 (SRR28752635), as found in \u2018avian-influenza\/reference\u2019, using mafft, then concatenated the segments longest-to-shortest into a single sequence per sample (assuming no appreciable\u00a0reassortment, which appears valid as of October 2025).<\/p>\n<p>For dating, we identified the 3,339 sequences with GenBank accessions providing day-resolved dates (SRR metadata dates typically only indicate year). We prepared these as input to Delphy; runs with the larger set that\u00a0includes those with uncertain dates exhibited severe convergence problems and were excluded. For geographical analysis, we used GenBank\u2019s \u2018geo_loc_name\u2019, which resolves to a US state in 3,194 of 3,383 accessions.<\/p>\n<p>Delphy was run twice for 20\u2009billion steps (trace every 2\u2009million steps, trees every 20\u2009million). We applied a Skygrid model with 22\u00a0month-long intervals ending at the latest tip (5 August 2025), with a log-space random walk prior with 3\u2009month halving\/doubling time (precision \u03c4\u2009=\u20096.15 when effective population sizes are in years). Coalescent prior discretization was increased to 1,000 cells to reduce discretization error.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-11012-6#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"For brevity, we summarize the essentials of Delphy\u2019s operation here; full details are provided in the Supplementary Information.&hellip;\n","protected":false},"author":2,"featured_media":774462,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[5589,59,102,4230,4231,34425,54140,90,4863,56,54,55],"class_list":["post-774461","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health","tag-epidemiology","tag-gb","tag-health","tag-humanities-and-social-sciences","tag-multidisciplinary","tag-phylogenetics","tag-phylogenomics","tag-science","tag-software","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/774461","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=774461"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/774461\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/774462"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=774461"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=774461"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=774461"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}