Curate reference tree: remove recombinant/frameshifted/low-quality sequences#465
Conversation
TestingTry in Nextclade Web: ScienceTaxonomy, genome, and reference [click to expand]Coxsackievirus A16 (CVA16) belongs to genus Enterovirus, species Enterovirus A, family Picornaviridae -- a positive-sense single-stranded RNA virus and one of the major causative agents of hand, foot, and mouth disease (HFMD) in children. The CVA16 genome is approximately 7,400 nt and encodes a single polyprotein processed by viral proteases into structural proteins (VP4, VP2, VP3, VP1 in P1) and non-structural proteins (2A-2C in P2, 3A-3D in P3). VP1 is the standard molecular typing target for enteroviruses (Oberste et al., J Clin Microbiol 1999; Chen et al., PLoS ONE 2013). The dataset uses a Static Inferred Ancestor (SIA) as reference rather than the historical G-10 prototype strain (U05876.1), which diverges substantially from circulating strains. The upstream workflow constructs this ancestor by outgroup-rooting a CVA16 phylogeny, reconstructing the ingroup MRCA, and filling alignment gaps from the prototype reference (inference workflow). Lineage nomenclature and recombination [click to expand]CVA16 genotypes are defined by VP1 phylogeny. Chen et al. (2013) established genotypes A (prototype G-10) and B, with B subdivided into B1a, B1b, and B1c. A recent synthesis recognizes genotypes A, B, and D, with B1/B2 under B and B1a/B1b/B1c under B1 (Han et al., Virus Evol 2024). Hassel et al. (2017) identified clade D with intertype recombinant origin. Whole-genome classifications differ because recombination changes phylogenetic relationships outside VP1. Han et al. identified recombinant forms RF-A through RF-E from the 3D region, while another whole-genome analysis proposed G-a through G-e (Chu et al., Heliyon 2024). Both B1a and B1b contain sequences from other Enterovirus A donors in the 5'UTR and P2/P3 regions (Chen et al., 2013). The dataset's approach of removing singleton recombinants while retaining recurring circulating recombinant forms is standard curation practice. The upstream workflow explicitly classifies excluded accessions as recombinant outliers, frameshifted sequences, or UTR-associated outliers (curation commit). Epidemiological context and tree coverage [click to expand]B1a, B1b, and B1c remain epidemiologically relevant. B1b dominated a long-term mainland China dataset (Han et al., Virus Evol 2020), while B1a predominated among Thai CVA16 sequences through 2022 (Noisumdaeng and Puthavathana, Sci Rep 2023). In Hangzhou, B1c rose sharply during 2024 (Xu et al., Front Microbiol 2025). The curated tree spans 1997-2025 with 741 tips. Year distribution shows strong representation in 2008-2024, with 2024 having the most tips (111). The 2020 dip (2 tips) reflects reduced surveillance during the COVID-19 pandemic. One 2025 tip (PZ117058) extends coverage to recent circulation. The dataset is maintained by ENPEN (European Non-Polio Enterovirus Network), with the workflow at enterovirus-phylo/nextclade_a16. All sequence data is from GenBank; no restricted sources are identified. Blocking issuesCorrectness concerns worth addressing before merge. 🔴 H1. Nucleotide mutation labels emptied without documentation [click to expand]The The generated Nextclade uses these maps to label private mutations in the UI and tabular output. Labeled substitutions receive distinct QC weights to flag potential contamination, co-infection, or recombination (mutation-label documentation, private-mutation algorithm). Effect: all former clade-associated private mutations become unlabeled and receive weight 1 instead of 1.5. User-visible mutation annotations disappear, QC scores change, and the index metadata becomes misleading. Neither the PR description nor the CHANGELOG mention this change. Fix: regenerate Non-blocking issuesConvention drift and minor inconsistencies. Fix if time allows. 🟡 M1. Broken JSON indentation in pathogen.json [click to expand]Several sections have inconsistent indentation compared to the base version [src]:
Effect: valid JSON, but confusing for future manual edits. Likely artifacts of manual editing after deleting the large Fix: reformat with 🟡 M2. Recombinant clade nomenclature lacks explicit mapping [click to expand]The README states that recombinant forms C-F are "also referred to as B2, B3, and D" [src]. The tree contains clades A, B1, B1a, B1b, B1c, C, D, E, F, and the parent label RFs. Published nomenclatures are not directly interchangeable:
The current text does not establish how dataset clades C-F map to these systems. Fix: add a compact mapping table giving each dataset clade, the genomic region used to define it, the corresponding published designation, and a defining citation. If C-F are dataset-specific labels, state that explicitly. 🟡 M3. No `meta.bugs` or `meta["source code"]` URLs in pathogen.json [click to expand]Neither Fix: add 🔵 L1. "Xu et al." citation not identifiable from CHANGELOG [click to expand]The CHANGELOG references "Xu et al." without a publication year or title [src]. The upstream workflow's testing file Fix: replace with "Xu et al. (2025)." 🔵 L2. Trailing whitespace in README [click to expand]
Fix: remove trailing space and rebuild generated output. Clade distribution9 clades, 783 -> 741 tips [click to expand]
All clades retained. Removals concentrated in B1a and B1b. NotesClick to expand
|
…on the inferred reference sequence Previous builds inferred the tree's root sequence independently of the dataset reference, causing mismatched mutations at the root and a Nextclade preprocessing error. The tree is now rebuilt with ancestral reconstruction anchored to the reference (inferred ancestral) sequence.
|
Thank you! My autocomplete machines are much happier now! The rest are nice to haves, but not a big deal.
Re-review of the updated branch:
TestingTry in Nextclade Web: Resolved since last review✅ H1 (was blocking).
|
| Clade | Base (master) |
Curated (7e9a9870) |
Removed | Added | Delta |
|---|---|---|---|---|---|
| A | 1 | 1 | |||
| B1 | 39 | 39 | |||
| B1a | 373 | 343 | |||
| B1b | 242 | 231 | |||
| B1c | 98 | 96 | |||
| C | 3 | 3 | |||
| D | 22 | 22 | |||
| E | 2 | 2 | |||
| F | 2 | 2 | |||
| unassigned | 1 | 1 | |||
| Total | 783 | 740 |
- All 9 named clades retained; removals concentrated in B1a and B1b.
- The single
unassignedtip is the SIA root leaf (ancestral_sequence), which carries no clade membership by design. - Curated total is 740, vs the prior review's 741, because commit
00ba506re-rooted the reconstruction and shifted one tip.
Notes
Click to expand
- N1. Reference representation (now clean).
ancestral_sequenceis a tree leaf at divergence 0.0 with 0 branch mutations; the prototypeU05876is also a leaf (div 0.243, 957 muts). No runtime-placement artifacts for either. This is the concrete payoff of the tree/reference re-rooting fix. - N2. Annotations. GFF3 defines 11 CDS (VP4->3D); all lengths are divisible by 3, and coordinates match
meta.genome_annotationsin tree.json exactly. Reference is 7,413 nt with zero ambiguous bases. - N3. Label map integrity.
nucMutLabelMapkeys are all well-formed (<pos><ALT>, positions 11-7413), no empty label lists, labels drawn from the clade set. Regenerated to match the curated topology. - N4. Example panel. 34 examples, lengths 592-7,410 nt; 21 are tree leaves, 13 placed at runtime. 6 now contain gap/
Ncharacters, added deliberately to exercise gap handling. - N5. Generated output. Source and
data_output/.../unreleased/copies are byte-identical for all text files exceptpathogen.json, which differs only by the rebuild-injectedversionblock (tag: unreleased) -- expected. Thedataset.zipmembers are SHA-256-identical to the unpacked files. - N6. Collection registration.
enpen/enterovirus/cva16is present indata/enpen/collection.jsonanddata_output/index.jsonregisters theunreleasedversion. - N7. Sibling comparison (EV-D68). cva16 uses
minSeedCover: 0.80vs ev-d68's 0.40, andnucMutLabelMap3,882 vs 2,057. Both ENPEN datasets carry an effectively emptyaaMutLabelMapand omitmetaURLs (see M3). - N8.
tree.jsonhas no trailing newline, matching the sibling ev-d68 tree -- consistent within the collection, not a regression. - N9. Cross-check with peer reviewer (Codex). An independent re-review reached the same conclusions: H1 resolved (3,882 labels, index consistent, reference at divergence 0), M2/M3/L1 remaining, clade table 783 -> 740.
add year to citation in CHANGELOG and add meta.bugs and source code to pathogen.json
|
Wow this is an amazing review bot - very cool Ivan!! |
Summary
Manually curated the dataset to improve reference tree quality: