Abstract
Phylogenetic placement, the problem of placing a 'query' sequence into a precomputed phylogenetic 'backbone' tree, is useful for constructing large trees, performing taxon identification of newly obtained sequences, and other applications. The most accurate current methods, such as pplacer and EPA-ng, are based on maximum likelihood and require that the query sequence be provided within a multiple sequence alignment that includes the leaf sequences in the backbone tree. This approach enables high accuracy but also makes these likelihood-based methods computationally intensive on large backbone trees, and can even lead to them failing when the backbone trees are very large (e.g., having 50,000 or more leaves). We present SCAMPP (SCaling AlignMent-based Phylogenetic Placement), a technique to extend the scalability of these likelihood-based placement methods to ultra-large backbone trees. We show that pplacer-SCAMPP and EPA-ng-SCAMPP both scale well to ultra-large backbone trees (even up to 200,000 leaves), with accuracy that improves on APPLES and APPLES-2, two recently developed fast phylogenetic placement methods that scale to ultra-large datasets. EPA-ng-SCAMPP and pplacer-SCAMPP are available at https://github.com/chry04/PLUSplacer.
Original language | English (US) |
---|---|
Pages (from-to) | 1417-1430 |
Number of pages | 14 |
Journal | IEEE/ACM Transactions on Computational Biology and Bioinformatics |
Volume | 20 |
Issue number | 2 |
DOIs | |
State | Published - Mar 1 2023 |
Externally published | Yes |
Keywords
- EPA-ng
- Phylogenetic placement
- maximum likelihood
- phylogenetics
- pplacer
ASJC Scopus subject areas
- Applied Mathematics
- Genetics
- Biotechnology
Fingerprint
Dive into the research topics of 'SCAMPP: Scaling Alignment-Based Phylogenetic Placement to Large Trees'. Together they form a unique fingerprint.Datasets
-
Biological and Simulated datasets for testing the SCAMPP framework for phylogenetic placement methods
Wedell, E. (Creator) & Warnow, T. (Creator), University of Illinois Urbana-Champaign, Apr 29 2022
DOI: 10.13012/B2IDB-9257957_V1
Dataset