Taxonomic classification accuracy on four materials of known composition, three nanopore chemistries and two competitors run against the same reference. Figures from the production engine; each one with the chemistry and the year of the material that produced it.
| Metric | Correct reads over classified reads, the literal definition given by GAIA's authors. Our coverage is published too, which they did not publish. |
| Rank | Genus in the comparison with the literature, because that is the rank Brown analysed. Species wherever the material supports it, with its denominator in plain sight. |
| Platforms | Nanopore: MinION, GridION and PromethION, chemistries R7.3/R9, R9.4.1 and R10.4.1. Illumina: not measured by this route. |
| Reference | RefSeq prokaryotic catalogue, 22,514 genomes. The results do not extend to fungi or viruses, and the yeasts in the standards fall outside the count for that reason. |
| Code | The bench imports the very classification function that runs in the product. Automatic test guards compare the published text against the run's JSON. |
A benchmark gets quoted in pieces. Species accuracy on Brown 2017 measures a 2015 chemistry with a median identity of 74% — at that data quality no tool on the market sustains a species call — not the tool. Each row declares the material that produced the number.
| Material | Year · chemistry | Identity | Genus | Species | Reads to species |
|---|---|---|---|---|---|
| Brown et al. 2017 · 7 MinION datasets | 2015 · R7.3/R9 | 74,23 % | 97,70 % | 37,4 % | 99 |
| ZymoBIOMICS D6300 · GridION | 2019 · R9.4.1 | 88,77 % | 99,08 % | 97,24 % | 15,549 |
| ZymoBIOMICS D6300 · PromethION | 2019 · R9.4.1 | 88,77 % | 99,00 % | 97,14 % | 14,731 |
| Zymo HMW mock · PRJNA934154 | 2023 · R10.4.1 | 97,30 % | 98,53 % | 97,26 % | 28.774 |
| Degraded series from D6300 | 2019 degraded | 74,23 % | 99,40 % | not claimed | — |
The better the chemistry, the more reads resolved to species at the same accuracy: 99 in 2015, 15,549 in 2019, 28,774 in 2023, always above 97% correct. Each accuracy is conditioned on the reads the system does classify — 89.5% in 2023 against 53.7% in 2019 — which is why the reads column travels next to the percentage in all five rows.
Brown et al. (2017) sequenced pure cultures and mixtures of declared composition on MinION and published the accuracy of MG-RAST, Kraken and One Codex. Paytuví-Gallart et al. (2019) reprocessed those same files with GAIA and added their column. It is the only point where an OmniOta figure can be placed beside a competitor's on known ground truth. The nine datasets were downloaded from the ENA and the read counts match the published ones exactly in all nine.
| Dataset | OmniOta | GAIA | Kraken | One Codex | MG-RAST |
|---|---|---|---|---|---|
| Ecoli · E. coli | 95,00 % | 100,0 % | 99,5 % | 98,7 % | 74,7 % |
| Pfluor · P. fluorescens | 94,64 % | 85,83 % | 84,6 % | 84,2 % | 84,9 % |
| Maeru · M. aeruginosa | 96,58 % | 96,39 % | 85,8 % | 95,1 % | 53,1 % |
| Selong · S. elongatus | 100,00 % | 100,0 % | 98,1 % | 97,6 % | 87,9 % |
| Equal (5) c | 97,70 % | 93,47 % | 97,6 % | 87,4 % | 65,0 % |
| Equal (6) c | 100,00 % | 98,94 % | 98,0 % | 98,7 % | 85,9 % |
| Rare (6) c | 100,00 % | 100,0 % | 99,1 % | 98,7 % | 92,9 % |
| Mean (7) | 97,70 % | 96,37 % | 94,67 % | 94,34 % | 77,77 % |
| Aggregate · 896/932 | 96,1 % | Wilson 95% CI [94.7% – 97.2%] — the interval contains GAIA's 96.37% | |||
c upper bound: in a mixture, any read assigned to any of the four members counts as correct. It affects all five tools equally, and their published figures were computed on those same mixtures. The GAIA, Kraken, One Codex and MG-RAST columns are the ones published by their authors: GAIA is a closed service with a proprietary database and has not been run, so part of the difference is attributable to the reference and not to the method.
The highest mean of the five tools is ours, and the largest advantage falls on the hardest dataset —P. fluorescens, where the four published tools land between 84.2% and 85.8%. A mean of seven percentages gives the same weight to a dataset of 27 reads as to one of 317; it is computed that way because that is what GAIA did and without it there would be no comparison, and the aggregate one — the statistically correct one — is published alongside it. At this data volume the mean advantage does not reach significance, and it is declared as such.
Shigella counts as Escherichia: it is polyphyletic within E. coli and its ANI exceeds the species threshold; the name is retained for clinical, not taxonomic, relevance (ISO 13136:2012). It affects 36 of the 78 reads in the Ecoli dataset — without the merge the figure would be 39.7% instead of 94.6%, and that is why it is declared. There are exactly two cases in the whole catalogue, they derive from the complexes the product already shows the user, and they do not touch species names.
The dataset staggered is HM-783D from BEI Resources: twenty certified species between 0.03% and 41.25%, with congeners that have to be separated —S. aureus de S. epidermidis, three Streptococcus— and members three orders below the dominant one. 96.86% at genus, 185 of 191 reads, an upper bound. Brown published 93% on this same material. It does not enter the mean because GAIA published no figure for it.
Ten species with composition certified in the manufacturer's datasheet — eight bacteria at 12% genomic DNA and two yeasts at 2% — sequenced by Nicholls, Quick, Tang and Loman (2019) on GridION and PromethION. 50,000 reads measured per platform.
| GridION | PromethION | |
|---|---|---|
| accuracy at genus | 99,08 % | 99,00 % |
| accuracy at species | 97,24 % | 97,14 % |
| reads to species | 15,549 | 14,731 |
| largest false positive | 0,17 % | 0,21 % |
| L1 between platforms | 0.0181 · expected from sampling (p95) 0.0257 · 0 divergent taxa → invariant | |
The composition is invariant between GridION and PromethION: the per-genus differences run from 0.05 to 0.43 points and the global divergence is statistically indistinguishable from sampling. What does not repeat are the low-abundance artefacts —Macrococcus, which is not a member, appears with 23 reads on one platform and 5 on the other. The composition repeats, the traces do not, and that settles how much weight to give a taxon held up by twenty reads. Accuracy is an upper bound because this is a mixture; the reads come from the run's prefix, which over-represents the flow cell's first hours, and it is stated as such.
A 2026 product evaluated on 2019 chemistry is a legitimate objection, and it is not answered with arguments but by measuring. Identity is measured by aligning against the standard's genomes; it is not taken from the specification sheet.
| Material | Chemistry | Identity |
|---|---|---|
| Brown · 2015 | R7.3 / R9 | 74,23 % |
| Nicholls · 2019 | R9.4.1 | 88,77 % |
| Zymo HMW · 2023 | R10.4.1 (sup) | 97,30 % |
On that chemistry the engine classifies 89.5% of the reads at genus with 98,53 % accuracy and 57.5% at species with 97,26 %, finding 7 of 7 members at both ranks. Against the 2019 material, reads resolved to species go from 31.1% to 57.5% at the same accuracy: nearly twice the resolution without losing precision.
The same run published with all three ONT basecallers — same signals, same pore, same day — gives fast 89,80 %, hac 96,16 % y sup 97,30 %. A 2023 run with fast basecalling lands at the same identity as the 2019 material with the best basecalling. Low identity is not a historical condition now overcome: it is a state entered today, through a configuration checkbox, through thermally degraded DNA, through a fatty matrix or through a flow cell at the end of its life. fast consumes 7.5 of the ~15 points of margin separating the modern material from the point where an exact k-mer method stops returning a result.
Between Brown 2017 and Zymo D6300 everything changes at once: year, chemistry, sequencer, community, laboratory and identity. With two points differing in six variables it is not justified to attribute a collapse to one of them. The experiment takes the same file from Zymo and injects error at increasing rates with a fixed seed: everything else is identical by construction. Substitutions only, no indels — which are more destructive for an exact k-mer — so the scenario is conservative in the competitor's favour.
| Identity | Members OmniOta | Taxa sylph |
|---|---|---|
| 88,77 % | 8/8 | 8 |
| 87,17 % | 8/8 | 7 |
| 84,91 % | 8/8 | 6 |
| 82,65 % | 8/8 | 0 |
| 79,68 % | 8/8 | 0 |
| 76,84 % | 8/8 | 0 |
| 74,23 % | 8/8 | 0 |
The most degraded level lands at 74.23%, the median identity of Brown 2017: no 2015 dataset was needed; degrading the 2019 one is enough. All eight organisms are recovered at every level, including the most severe, where the competitor does not return a single line. It is the result that matters for a real food sample: DNA fragmented by processing, a fatty matrix, difficult extraction. Genus accuracy rises as the material degrades (99.07 → 99.40%) and that is not a merit but selection bias: the easy reads are what remain, and classified reads fall from 53.7% to 36.6%. On modern chemistry sylph works well: it detects 7 of 7 members on all three basecallers, including fast.
Against GAIA only its published figures are available: closed service, proprietary database. With sylph the reference can indeed be equalised, and it has been, genome by genome, on certified D6300. Both tools find the complete community and each is stronger on a different axis.
| OmniOta | sylph | |
|---|---|---|
| per-read accuracy | 99,18 % | not applicable |
| members found | 8/8 | 8/8 |
| taxa emitted | 61 | 8 |
| weight of the surplus | 0,82 % | — |
Per read OmniOta is highly precise, and that is the figure governing an abundance table and everything computed on top of it. Per named taxon sylph decides better: it models genome coverage with k-mer statistics and requires real presence before naming, whereas we name any taxon that receives an assignment. The criterion under development is a minimum evidence to name, not merely to assign: the same Wilson doctrine the product already applies to per-taxon confidence, carried into the decision to publish.
On the HMW standard, equimolar at 14.3%, counting reads gave Salmonella el 47,1 % y a Bacillus 3.2%. Abundance is mean coverage — aligned bases over genome size — and with that metric the L1 divergence falls from 0.852 to 0,411: measuring DNA mass halves the error. MetaPhlAn 4.2.6, which shares nothing with our method — UniRef markers against complete genomes — agrees with us to within 2.14 points on average and on the ordering of the three largest: the sequenced material is not equimolar, and that is established by two independent routes. On what it does compete on, OmniOta lands 18% closer to the certified composition (L1 0.411 against 0.501). MetaPhlAn detects S. cerevisiae and we do not, because our reference is bacterial.
| version | 1.4.0 |
| date of measurement | 2026-08-10 |
| engine | minimap2-lca/3.0.0 |
| working tree | 61a0bb7c |
| reference | 22,514 genomes · 35.1 GB |
| competitors here | sylph · MetaPhlAn 4.2.6 |
Kraken2/Bracken is not measured: its database with these same 22,514 genomes would require on the order of 170 GB of RAM, as measured, so the reason is published instead of giving it a trimmed database.