{"id":248167,"date":"2025-11-06T22:31:20","date_gmt":"2025-11-06T22:31:20","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/248167\/"},"modified":"2025-11-06T22:31:20","modified_gmt":"2025-11-06T22:31:20","slug":"assessing-phylogenetic-confidence-at-pandemic-scales","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/248167\/","title":{"rendered":"Assessing phylogenetic confidence at pandemic scales"},"content":{"rendered":"<p>SPRTA and aBayes<\/p>\n<p>Given an estimated rooted phylogenetic tree T and data D in the form of a multiple sequence alignment, our aim is to assign confidence scores to branches b of T.<\/p>\n<p>We take inspiration from the approximate Bayes (aBayes) approach<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Anisimova, M., Gil, M., Dufayard, J.-F., Dessimoz, C. &amp; Gascuel, O. Survey of branch support methods demonstrates accuracy, power, and robustness of fast likelihood-based approximation schemes. Syst. Biol. 60, 685&#x2013;699 (2011).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR17\" id=\"ref-link-section-d37827382e2023\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>. aBayes assigns to b a probability score Pr(b\u2223D,\u00a0T\\b) based on the ratio of the likelihood Pr(D\u2223T) of the estimated binary tree T versus the likelihoods of the trees \\({T}_{i}^{b}\\) obtained using nearest neighbour interchange<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Swofford, D., Olsen, G., Waddell, P. &amp; Hillis, D. in Molecular Systematics (eds Hillis, D. M. et al.) 407&#x2013;514 (Sinauer Associates, 1996).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR36\" id=\"ref-link-section-d37827382e2085\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a> (NNI) moves centred around b (which comprise \\(T={T}_{1}^{b}\\) itself in addition to two tree topologies not containing the clade in T descending from b):<\/p>\n<p>$${\\rm{aBayes}}(b)=\\Pr (b| D,T\\backslash b)=\\frac{\\Pr (D| T)}{{\\sum }_{i=1}^{3}\\Pr (D| {T}_{i}^{b})}.$$<\/p>\n<p>\n                    (2)\n                <\/p>\n<p>These NNI moves perform small changes to T adjacent to branch b. One of the appeals of aBayes is that it can score not only T, but also the two alternative topologies considered, \\({T}_{2}^{b}\\) and \\({T}_{3}^{b}\\), not containing the clade defined by b:<\/p>\n<p>$$\\Pr ({T}_{j}^{b}| D,T\\backslash b)=\\frac{\\Pr (D| {T}_{j}^{b})}{{\\sum }_{i=1}^{3}\\Pr (D| {T}_{i}^{b})}.$$<\/p>\n<p>\n                    (3)\n                <\/p>\n<p>aBayes has a topological focus, with score aBayes(b) for branch b interpreted as the support for the existence of the clade of T containing all descendants of b. aBayes(b) is in effect an approximate Bayesian posterior score for this clade, where a flat tree prior is assumed. In contrast to typical Bayesian phylogenetics, however, instead of integrating over branch lengths we define \\(\\Pr (D| {T}_{i}^{b})\\) as the maximum-likelihood score of topology \\({T}_{i}^{b}\\) over all possible branch lengths, and we only consider the alternative topologies obtainable through a single NNI move<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Anisimova, M., Gil, M., Dufayard, J.-F., Dessimoz, C. &amp; Gascuel, O. Survey of branch support methods demonstrates accuracy, power, and robustness of fast likelihood-based approximation schemes. Syst. Biol. 60, 685&#x2013;699 (2011).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR17\" id=\"ref-link-section-d37827382e2639\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a> starting from T. Although computationally much faster than Felsenstein\u2019s bootstrap, aBayes can still be too demanding for large genomic epidemiological datasets if implemented within classical maximum-likelihood phylogenetic methods (for example, see Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). Also, aBayes is defined based only on NNI moves, an approach insufficiently comprehensive for pandemic-scale data<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e2649\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>: owing to the existence of many phylogenetic topologies with similar likelihood<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 29\" title=\"Morel, B. et al. Phylogenetic analysis of SARS-CoV-2 data is difficult. Mol. Biol. Evol. 38, 1777&#x2013;1791 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR29\" id=\"ref-link-section-d37827382e2654\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a>, the two alternative topologies obtainable through NNI moves, \\({T}_{2}^{b}\\) and \\({T}_{3}^{b}\\), might represent only a very small subset of plausible alternative topologies not containing the clade defined by b.<\/p>\n<p>Here we address these limitations and define a new measure of branch support, SPRTA, that is particularly useful in the context of large-scale genomic epidemiology, but is also applicable more generally in phylogenetics. First, to address the problem of computational demand, we consider trees estimated using methods suitable for pandemic-scale datasets, such as MAPLE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e2730\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>. In the following, we assume that the likelihood of alternative tree topologies is also calculated with MAPLE or a similar method.<\/p>\n<p>We define the SPRTA support of branch b by considering a certain number Ib of possible alternative topologies \\({T}_{i}^{b}\\) (1\u2009\u2a7d\u2009i\u2009\u2a7d\u2009Ib), obtained by performing single subtree prune and regraft<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Swofford, D., Olsen, G., Waddell, P. &amp; Hillis, D. in Molecular Systematics (eds Hillis, D. M. et al.) 407&#x2013;514 (Sinauer Associates, 1996).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR36\" id=\"ref-link-section-d37827382e2790\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a> (SPR) moves that relocate Sb, the subtree of T containing all descendants of b, as a descendant of other parts of T (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). Again, we assume for simplicity that \\({T}_{1}^{b}=T\\) corresponds to the null SPR move. Compared to NNI moves, SPR moves are much more numerous, and can cause long-range changes to the tree; as such, SPR moves create a much more comprehensive set of alternative evolutionary histories than NNI moves. For each branch b, the number of possible SPR moves is linear in the number of sequences in the considered dataset, which means if we do not select which of these SPR moves we focus on, Ib can become too large, and the calculation of SPRTA scores too computationally demanding. However, to obtain an accurate evaluation we only need to consider topologies that have non-negligible likelihood scores compared to T.<\/p>\n<p>For this reason, we first perform an initial, approximate evaluation of alternative tree topologies obtained through SPR relocations of Sb using fixed branch lengths (we leave the length of the placement branch equal to the length of b, and we assess placements only halfway along branches). From this, we retain only topologies with an initial log-likelihood score difference from T within a threshold corresponding approximately to one extra mutation event (the natural logarithm of the genome length, which for SARS-CoV-2 is about 10.3). The likelihoods of all Ib topologies passing this threshold are then more deeply evaluated by optimizing branch lengths. This two-step approach is similar to the \u2018baseball\u2019 heuristic of the metagenomic query mapper pplacer<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Matsen, F. A., Kodner, R. B. &amp; Armbrust, E. V. pplacer: linear time maximum-likelihood and bayesian phylogenetic placement of sequences onto a fixed reference tree. BMC Bioinformatics 11, 538 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR19\" id=\"ref-link-section-d37827382e2888\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a>. More precisely, when we consider an alternative placement of Sb on branch \\({b}^{^{\\prime} }\\), let us use \\({n}^{^{\\prime} }\\) to denote the new node at the conjunction of \\({b}^{^{\\prime} }\\) with Sb. For this placement, the three branches whose length we optimize are the one connecting \\({n}^{^{\\prime} }\\) with Sb (length l1), the one within \\({b}^{^{\\prime} }\\) below \\({n}^{^{\\prime} }\\) (length l2), and the one within \\({b}^{^{\\prime} }\\) above \\({n}^{^{\\prime} }\\) (length l3), similarly to the \u2018lazy subtree rearrangement\u2019 approach of RAxML<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Stamatakis, A., Ludwig, T. &amp; Meier, H. RAxML-III: a fast program for maximum likelihood-based inference of large phylogenetic trees. Bioinformatics 21, 456&#x2013;463 (2005).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR37\" id=\"ref-link-section-d37827382e3149\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a> (see Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10a<\/a>). \\(\\Pr (D| {T}_{i}^{b})\\) is defined as the maximum likelihood obtained by optimizing these three branches, while leaving all other model parameters and branch lengths unaltered. This approximation of \\(\\Pr (D| {T}_{i}^{b})\\) is the same one made by MAPLE, and is justified by the fact that, at the low levels of divergence typically considered in genomic epidemiology, changes in topology and branch lengths usually only affect nodes near those directly affected by the changes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e3261\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>.<\/p>\n<p>The likelihood score \\(\\Pr (D| {T}_{i}^{b})\\) of an SPR move is calculated by MAPLE (up to a normalizing factor) by comparing the partial likelihoods at node B (the root of Sb), informed by sequence data within Sb, against the partial likelihoods at the new placement nodes Ai, informed by the sequence data within T\\Sb (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e3351\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>). The partial likelihoods at node B represent genetic sequence uncertainty for the ancestor represented by this node. In genomic epidemiology, due to dense sampling and short divergence, this uncertainty is often negligible, and these partial likelihoods will often support only one genome. Likewise, partial likelihoods at Ai will often also only support a single genome. In these cases, the score \\(\\Pr (D| {T}_{i}^{b})\\) (up to a normalizing factor) is used as an approximation of the probability that the genome in B evolved from the genome in Ai. In case uncertainty of the ancestral genetic sequence at these nodes is not negligible, we consider all possible ancestral genomes, weighted according to their likelihood, and \\(\\Pr (D| {T}_{i}^{b})\\) will approximate the probability that any of the possible genomes in B evolved from any of the possible genomes in Ai.<\/p>\n<p>Due to the initial filtering of plausible alternative topologies, Ib depends on branch b. These topologies are the same ones typically considered during the final stage of standard tree search in MAPLE.<\/p>\n<p>SPRTA defines the support probability of b as<\/p>\n<p>$${\\rm{SPRTA}}(b)=\\Pr (b| D,T\\backslash b)=\\frac{\\Pr (D| T)}{\\sum _{1\\leqslant i\\leqslant {I}_{b}}\\Pr (D| {T}_{i}^{b})}$$<\/p>\n<p>\n                    (4)\n                <\/p>\n<p>where T\u2009\\\u2009b represents the two subtrees obtained by removing b from T. We can similarly define \\(\\Pr ({T}_{i}^{b}| D,T\\backslash b)\\) for the alternative placements \\({T}_{i}^{b}\\ne T\\) (2 \u2a7d i \u2a7d Ib) for any subtree Sb considered. Note that SPRTA typically evaluates a much larger number of alternative topologies than aBayes (equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Equ2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>)): up to quadratic in the number of genomes analysed for SPRTA, but only linear for other local support measures such as aBayes or aLRT.<\/p>\n<p>When evaluating possible SPR placements, placements resulting in the same topology are equivalent, and so we only consider them once. In particular, we take into account multifurcations caused by branches of length 0, and consider placements into points of the tree with a null distance between them (based on considering optimized placement-specific branch lengths as just described) as equivalent.<\/p>\n<p>Because relative likelihood calculations of alternative topologies are performed as a standard component of tree inference in MAPLE, assessing SPRTA support probabilities adds negligible computational overhead.<\/p>\n<p>MAPLE infers rooted phylogenetic trees, and as such we have defined and implemented SPRTA to assess rooted tree inference. The same principles could however also be applied to the evaluation of unrooted trees. In this case, to calculate the score of a branch b, we would not only consider SPR moves representing alternative placements of subtree Sb within T\u2009\\\u2009Sb, but also SPR moves representing alternative placements of T\u2009\\\u2009Sb within Sb, so increasing the number of SPR moves to be considered for each branch b, but leaving the rest of the approach unaltered.<\/p>\n<p>Benchmarking of SPRTA support<\/p>\n<p>Branch support measures are often used to assess the expected accuracy of phylogenetic inference. However, how do we assess the accuracy of these measures of accuracy? We approach this problem using simulations, for which we have a ground truth against which to compare estimates. First, we simulate a tree and a set of genomes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>); second, we estimate a tree from these simulated genomes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10c<\/a>). Some branches of the inferred tree will be correct and some will be wrong: therefore, third, we assess branch support measures according to their ability to give higher support scores to correctly estimated branches, and lower support scores to wrongly estimated branches (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10d<\/a>).<\/p>\n<p>We benchmark branch support methods on simulated SARS-CoV-2 genome data (Methods, &#8216;Simulated genomes\u2019) using our \u2018mutational\u2019 focus rather than a traditional \u2018topological\u2019 focus. We define phylogenetic correctness in terms of the genome evolutionary history implied by a phylogenetic tree. From each individual simulated dataset (see graphical example in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>), we first infer in MAPLE a maximum-likelihood phylogenetic tree T and mutation events by marginal posterior mutation mapping<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 38\" title=\"Nielsen, R. Mapping mutations on phylogenies. Syst. Biol. 51, 729&#x2013;739 (2002).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR38\" id=\"ref-link-section-d37827382e3872\" rel=\"nofollow noopener\" target=\"_blank\">38<\/a> conditional on T (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10c<\/a>). Only mutation events inferred with \u00a0&gt;0.5 probability by MAPLE are considered here. We define estimated mutation events as pairs (G,\u00a0m) where G is a whole genome sequence and m\u2009=\u2009(n1,\u00a0p,\u00a0n2) is a single-nucleotide substitution at position p of G from nucleotide n1 (contained in G at position p) to nucleotide n2 (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10d<\/a>). An inferred mutation event (G,\u00a0m) is considered correct if it is also present in the simulated tree, otherwise it is considered as an estimation error (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10d<\/a>). Finally, we assign the support score of branch b (estimated by SPRTA or any other method) to all the mutations inferred to have occurred on b. We consider a branch support measure as more accurate if in our simulations it assigns higher support to correctly estimated mutation events, and lower support to erroneously inferred ones.<\/p>\n<p>Evaluating the accuracy of SPRTA using inferred mutations might seem counter-intuitive, since our definition of SPRTA scores does not consider explicit mutational histories. However, by analysing the levels of support for different placements of subtree Sb associated with branch b, different possible mutation histories immediately ancestral to Sb are implicitly considered and evaluated via the likelihood of alternative subtree placements. We thus interpret SPRTA scores as the support for the hypothesis that the profile B at the lower end of b evolved from profile A at the upper end of b through mutations on branch b (see \u2018Accuracy\u2019).<\/p>\n<p>In the scenario of short branches considered here, there is extremely low uncertainty in most mutation events and ancestral genomes implied by a given topology, and so we can often interpret the correctness of the mutational history inferred on b as the correctness of b itself. This does not mean that in this case b will be topologically correct\u2014for example, the placement of rogue taxa within or outside the subtree Sb defined by b can make b topologically uncertain without causing uncertainty in the inferred mutational history or the placement of Sb.<\/p>\n<p>Implementation and usage of SPRTA<\/p>\n<p>We ran SPRTA as implemented within MAPLE v.0.6.8 (<a href=\"https:\/\/github.com\/NicolaDM\/MAPLE\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/NicolaDM\/MAPLE<\/a>). While SPRTA can be run in MAPLE at the same time as tree inference with minimal additional computational cost (Methods \u2018SPRTA and aBayes\u2019), to aid comparability of computational performance with other approaches here we have considered its use to assess a pre-estimated input phylogenetic tree provided with the option \u2013inputTree. We used options \u2013numTopologyImprovements 0 \u2013doNotImproveTopology to perform a shallow SPR search in MAPLE. We also used options \u2013model UNREST \u2013rateVariation to use an UNREST model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 39\" title=\"Yang, Z. Estimating the pattern of nucleotide substitution. J. Mol. Evol. 39, 105&#x2013;111 (1994).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR39\" id=\"ref-link-section-d37827382e4027\" rel=\"nofollow noopener\" target=\"_blank\">39<\/a> with rate variation<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"De Maio, N. et al. Rate variation and recurrent sequence errors in pandemic-scale phylogenetics. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2024.07.12.603240&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR21\" id=\"ref-link-section-d37827382e4031\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>, and option \u2013estimateMAT to infer mutation events.<\/p>\n<p>Other branch support methods<\/p>\n<p>All other branch support measures considered here were calculated using IQ-TREE v.2.1.3<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 40\" title=\"Minh, B. Q. et al. IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Mol. Biol. Evol. 37, 1530&#x2013;1534 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR40\" id=\"ref-link-section-d37827382e4043\" rel=\"nofollow noopener\" target=\"_blank\">40<\/a> with options \u2013seqtype DNA \u2013seed 1 -m GTR+F+G4 \u2013quiet -nt 1. As with SPRTA, we always use the tree estimated by MAPLE as a starting tree (supplied via the option -t) since on these datasets IQ-TREE will typically not converge to a tree with likelihood as high as MAPLE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e4047\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>. We used the following additional IQ-TREE options:<\/p>\n<p>-B 1000 (1,000 bootstrap replicates) for UFBoot2<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Hoang, D. T., Chernomor, O., von Haeseler, A., Minh, B. Q. &amp; Vinh, L. S. Ufboot2: improving the ultrafast bootstrap approximation. Mol. Biol. Evol. 35, 518&#x2013;522 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR7\" id=\"ref-link-section-d37827382e4057\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a><\/p>\n<p>\u2013fast -b 100 (100 bootstrap replicates and fast tree search) for Felsenstein\u2019s bootstrap<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Felsenstein, J. Confidence limits on phylogenies: an approach using the bootstrap. Evolution 39, 783&#x2013;791 (1985).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR2\" id=\"ref-link-section-d37827382e4066\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a><\/p>\n<p>\u2013fast -b 100 \u2013tbe (100 bootstrap replicates and fast tree search) for TBE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Lemoine, F. et al. Renewing Felsenstein&#x2019;s phylogenetic bootstrap in the era of big data. Nature 556, 452&#x2013;456 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR10\" id=\"ref-link-section-d37827382e4075\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a><\/p>\n<p>\u2013fast \u2013alrt 1000 (1,000 bootstrap replicates and fast tree search) for aLRT-SH<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Guindon, S. et al. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of phyml 3.0. Syst. Biol. 59, 307&#x2013;321 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR16\" id=\"ref-link-section-d37827382e4084\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a><\/p>\n<p>\u2013fast \u2013alrt 0 (fast tree search) for aLRT<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Anisimova, M. &amp; Gascuel, O. Approximate likelihood-ratio test for branches: a fast, accurate, and powerful alternative. Syst. Biol. 55, 539&#x2013;552 (2006).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR15\" id=\"ref-link-section-d37827382e4093\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a><\/p>\n<p>\u2013fast \u2013abayes (fast tree search) for aBayes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Anisimova, M., Gil, M., Dufayard, J.-F., Dessimoz, C. &amp; Gascuel, O. Survey of branch support methods demonstrates accuracy, power, and robustness of fast likelihood-based approximation schemes. Syst. Biol. 60, 685&#x2013;699 (2011).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR17\" id=\"ref-link-section-d37827382e4103\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a><\/p>\n<p>\u2013fast \u2013lbp 1000 (1,000 bootstrap replicates and fast tree search) for LBP<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"Adachi, J. &amp; Hasegawa, M. MOLPHY version 2.3: programs for molecular phylogenetics based on maximum likelihood (Institute of Statistical Mathematics, 1996).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR14\" id=\"ref-link-section-d37827382e4112\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>.<\/p>\n<p>The number of bootstrap replicates has very limited impact on the computational demand of UFBoot2 and LBP, hence these were set to 1,000 for reduced stochasticity with minimal computational cost.<\/p>\n<p>SARS-CoV-2 genome datasetsViridian genome dataset<\/p>\n<p>We applied SPRTA to a SARS-CoV-2 dataset containing 2,072,111 genomes collected up to February 2023. The consensus sequences of these genomes were consistently called with Viridian, a tool that prevents common reference biases in genomic regions of low sequencing coverage<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 23\" title=\"Hunt, M. et al. Addressing pandemic-wide systematic errors in the SARS-CoV-2 phylogeny. Nat. Methods (in the press).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR23\" id=\"ref-link-section-d37827382e4132\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a>. Furthermore, we filtered out potentially contaminated samples, and masked alignment columns affected by recurrent sequence errors<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"De Maio, N. et al. Rate variation and recurrent sequence errors in pandemic-scale phylogenetics. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2024.07.12.603240&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR21\" id=\"ref-link-section-d37827382e4136\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>. We estimated a phylogenetic tree using MAPLE v.0.6.8 with an UNREST substitution model, rate variation, and deep SPR phylogenetic search. For a full description of data preparation and phylogenetic inference, see ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"De Maio, N. et al. Rate variation and recurrent sequence errors in pandemic-scale phylogenetics. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2024.07.12.603240&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR21\" id=\"ref-link-section-d37827382e4140\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>. We ran SPRTA on this alignment and tree using MAPLE v.0.6.9, with the options described in \u2018Implementation and usage of SPRTA\u2019, and additionally with option \u2013supportFor0Branches to evaluate the support of all genome placements, even those not involving mutations (see \u2018Uncertainty of SARS-CoV-2 evolution\u2019).<\/p>\n<p>Simulated genomes<\/p>\n<p>For benchmarking, we simulated SARS-CoV-2 genomes evolving along a known (\u2018true\u2019) background phylogeny. The background tree we used was the publicly available 26 October 2021 global SARS-CoV-2 phylogenetic tree from <a href=\"http:\/\/hgdownload.soe.ucsc.edu\/goldenPath\/wuhCor1\/UShER_SARS-CoV-2\/\" rel=\"nofollow noopener\" target=\"_blank\">http:\/\/hgdownload.soe.ucsc.edu\/goldenPath\/wuhCor1\/UShER_SARS-CoV-2\/<\/a><a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 41\" title=\"McBroome, J. et al. A daily-updated database and tools for comprehensive SARS-CoV-2 mutation-annotated trees. Mol. Biol. Evol. 38, 5819&#x2013;5824 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR41\" id=\"ref-link-section-d37827382e4158\" rel=\"nofollow noopener\" target=\"_blank\">41<\/a> representing the evolutionary relationship of 2,250,054 SARS-CoV-2 genomes, as inferred using UShER<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Turakhia, Y. et al. Ultrafast sample placement on existing trees (usher) enables real-time phylogenetics for the SARS-CoV-2 pandemic. Nat. Genet. 53, 809&#x2013;816 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR8\" id=\"ref-link-section-d37827382e4162\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>.<\/p>\n<p>We used phastSim v.0.0.3<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"De Maio, N. et al. phastSim: efficient simulation of sequence evolution for pandemic-scale datasets. PLoS Comput. Biol. 18, e1010056 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR42\" id=\"ref-link-section-d37827382e4169\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a> with options \u2013treeFile public-latest.all.nwk \u2013scale 0.00003344 \u2013reference MN908947.3.fasta \u2013alpha 0.2 \u2013createNewick to simulate sequence evolution along this tree according to SARS-CoV-2 non-stationary neutral mutation rates<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"De Maio, N. et al. Mutation rates and selection on synonymous mutations in SARS-CoV-2. Genome Biol. Evol. 13, evab087 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR22\" id=\"ref-link-section-d37827382e4173\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>, using the SARS-CoV-2 Wuhan-Hu-1 genome<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 43\" title=\"Wu, F. et al. A new coronavirus associated with human respiratory disease in China. Nature 579, 265&#x2013;269 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR43\" id=\"ref-link-section-d37827382e4177\" rel=\"nofollow noopener\" target=\"_blank\">43<\/a> as root sequence, and with gamma-distributed (\u03b1\u2009=\u20090.2) rate variation<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Yang, Z. Among-site rate variation and its impact on phylogenetic analyses. Trends Ecol. Evol. 11, 367&#x2013;372 (1996).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR44\" id=\"ref-link-section-d37827382e4184\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a> (similar to that estimated from real data<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"De Maio, N. et al. Rate variation and recurrent sequence errors in pandemic-scale phylogenetics. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2024.07.12.603240&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR21\" id=\"ref-link-section-d37827382e4189\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>). These simulations of complete genomes were used for Figs. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>, and Extended Data Figs. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">1e<\/a>, <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>.<\/p>\n<p>We also created a second set of simulations mimicking the distribution of genome incompleteness from real data as in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e4211\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>. In each simulated sequence we included N and gap (\u2013) characters copied in number and location from a randomly selected paired sequence from the real SARS-CoV-2 genome dataset considered in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e4215\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>. This step simulates the distribution of missing sequence data due to low sequencing depth at certain specific genome regions. Additionally, in each simulated sequence we masked a number of randomly selected SNPs (differences with respect to the reference genome) equal in number to the isolated ambiguous characters in the paired randomly sampled real sequence. This additional step mimics the pattern caused by mixed infections and contamination, in which phylogenetically informative positions are selectively masked in consensus genomes due to within-sample heterozygosity (see ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"De Maio, N. et al. Maximum likelihood pandemic-scale phylogenetics. Nat. Genet. 55, 746&#x2013;752 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR9\" id=\"ref-link-section-d37827382e4219\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a> for more detail). This second set of simulations was used to create Extended Data Figs. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">1a\u2013d,f<\/a>, <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig10\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>.<\/p>\n<p>Assessing the impact on mutation rates<\/p>\n<p>To assess the impact of phylogenetic uncertainty on estimates of mutation patterns, we mimicked typical studies of mutation rate inference in SARS-CoV-2<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"De Maio, N. et al. Mutation rates and selection on synonymous mutations in SARS-CoV-2. Genome Biol. Evol. 13, evab087 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR22\" id=\"ref-link-section-d37827382e4242\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 45\" title=\"Simmonds, P. Rampant C&#x2009;&#x2192;&#x2009;U hypermutation in the genomes of SARS-CoV-2 and other coronaviruses: causes and consequences for their short- and long-term evolutionary trajectories. mSphere 5, e00408-20 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR45\" id=\"ref-link-section-d37827382e4245\" rel=\"nofollow noopener\" target=\"_blank\">45<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 46\" title=\"Haddox, H. K. et al. The mutation rate of SARS-CoV-2 is highly variable between sites and is influenced by sequence context, genomic region, and RNA structure. Nucleic Acids Res. 53, gkaf503 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR46\" id=\"ref-link-section-d37827382e4248\" rel=\"nofollow noopener\" target=\"_blank\">46<\/a>. We consider real SARS-CoV-2 data and the corresponding inferred tree and mutation events as described in \u2018Viridian genome dataset\u2019. We then created three datasets: one containing all inferred mutations, one containing only mutations on branches with at least 50% SPRTA support (representing a mildly conservative approach, discarding highly uncertain mutations), and one with only mutations on branches with at least 90% SPRTA support (representing a highly conservative approach, removing any moderately uncertain mutation).<\/p>\n<p>Equilibrium frequencies (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">6a<\/a>) represent the equilibrium nucleotide distribution of the Markov chain defined by the genome-wide mutation rate matrix<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Li&#xF2;, P. &amp; Goldman, N. Models of molecular evolution and phylogeny. Genome Res. 8, 1233&#x2013;1244 (1998).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR47\" id=\"ref-link-section-d37827382e4258\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a> inferred from the mutation counts as<\/p>\n<p>$${q}_{ij}\\propto \\frac{{n}_{ij}}{{{\\rm{\\pi }}}_{i}},$$<\/p>\n<p>\n                    (5)\n                <\/p>\n<p>where nij is the count of mutations from nucleotide i to nucleotide j and \u03c0i is the frequency of nucleotide i in the reference genome.<\/p>\n<p>We assess the impact of phylogenetic uncertainty on site-specific mutation counts by calculating, for every genome position, the ratio of the most conservative mutation counts (those on branches with at least 90% support) to the least conservative mutation counts (those on all branches) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">6b<\/a>). To avoid high variance in values of this ratio at sites with low numbers of substitutions, only sites with at least 50 mutations of the given type over all branches were included.<\/p>\n<p>Assessing the impact on Pango lineages<\/p>\n<p>To assess the impact of phylogenetic uncertainty on the definition and the inference of the origin of Pango lineages, we mapped 1,542 Pango lineage consensus genomes (as of 28 February 2023; <a href=\"https:\/\/github.com\/corneliusroemer\/pango-sequences\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/corneliusroemer\/pango-sequences<\/a>) onto our SARS-CoV-2 phylogenetic tree (\u2018Viridian genome dataset\u2019) using MAPLE v.0.7.3. This mapping associates Pango lineages with nodes in our tree. Similarly to our SARS-CoV-2 alignment, we masked regions of recurrent sequence errors from these consensus genomes (see ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"De Maio, N. et al. Rate variation and recurrent sequence errors in pandemic-scale phylogenetics. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2024.07.12.603240&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#ref-CR21\" id=\"ref-link-section-d37827382e4382\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>). Of these genomes, 1,127 mapped onto our tree with \u22641 mutation separating them from the tree, and with \u226595% SPRTA placement support score. We discarded consensus genomes that were further removed (&gt;1 mutation separating them from the tree), as they typically represent more-recent lineages that did not exist at the time our dataset was collected. Consensus genomes with low SPRTA placement support were discarded since they could not be uniquely associated with a single node in our tree (this uncertainty can be caused, for example, by incomplete genome sequences in our dataset). In total, these 1,127 genomes mapped onto 1,117 distinct tree branches, and represent the Pango lineages that we can place on the tree with high confidence.<\/p>\n<p>We used these 1,127 consensus genome placements to assign a Pango lineage to each branch and sample in our tree: each was assigned to the lineage represented by the closest ancestral consensus genome placement (the first one met moving from the considered node towards the tree root). We used this assignment of Pango lineages to assess uncertainty in the lineage assignment of the genomes in our dataset. For each of the 2,072,111 genomes in our alignment, we considered the SPRTA scores of their current and alternative placements. To each considered placement, we assigned the Pango lineage of the placement branch (alternative placements on the same branch as the placement of a Pango consensus genome were ignored).<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09567-x#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"SPRTA and aBayes Given an estimated rooted phylogenetic tree T and data D in the form of a&hellip;\n","protected":false},"author":2,"featured_media":248168,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[103665,59,102,4230,9577,4231,34425,34577,90,19174,56,54,55],"class_list":["post-248167","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health","tag-classification-and-taxonomy","tag-gb","tag-health","tag-humanities-and-social-sciences","tag-molecular-evolution","tag-multidisciplinary","tag-phylogenetics","tag-phylogeny","tag-science","tag-statistical-methods","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/248167","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=248167"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/248167\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/248168"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=248167"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=248167"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=248167"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}