Taxonomic classification accuracy measured on four materials of known composition, three nanopore chemistries and two competitors run against the same reference. Figures from the production engine, not from a laboratory variant.
Every figure in this dossier is measured on material of known composition, with the metric the competitor defined, and is published alongside the chemistry and the year of the material that produced it. Without that pair the figure describes a tool that does not exist: the same engine resolves 99 reads to species on 2015 chemistry and 28,774 on 2023 chemistry.
| Metric | Correct reads over classified reads, the literal definition given by GAIA's authors. Our coverage is published too, which they did not publish. |
| Rank | Genus in the comparison with the literature, because that is the rank Brown analysed. Species wherever the material supports it, with its denominator in plain sight. |
| Platforms | Nanopore: MinION, GridION and PromethION, chemistries R7.3/R9, R9.4.1 and R10.4.1. Illumina: not measured by this route. |
| Reference | RefSeq prokaryotic catalogue, 22,514 genomes. The results do not extend to fungi or viruses, and the two yeasts in the standards fall outside the count for that reason. |
| Code | The bench imports the very classification function that runs in the product, stamp minimap2-lca/3.0.0. Automatic test guards pin every methodological decision in this document. |
Document derived from the canonical documentation in the repository: BENCHMARK_OMNIOTA_vs_GAIA_BROWN2017.md v1.4.0, BANCO_ZYMO_D6300_INVARIANCIA.md, QUIMICA_ACTUAL_R1041.md and CURVA_DE_ROTURA_IDENTIDAD.md. No figure has been computed in this document: all of them are measured, versioned and reproducible from the repository.
It is the first table in the document for an operational reason: a benchmark gets quoted in pieces. Species accuracy on Brown 2017 measures a 2015 chemistry with a median identity of 74%, not the tool — and no tool on the market sustains a species call at that data quality. Each row declares the material that produced the number.
| Material | Year · chemistry | Identity | Genus | Species | Reads to species |
|---|---|---|---|---|---|
| Brown et al. 2017 · 7 MinION datasets | 2015 · R7.3/R9 | 74,23 % | 97,70 % | 37,4 % | 99 |
| ZymoBIOMICS D6300 · GridION composition certified by the manufacturer |
2019 · R9.4.1 | 88,77 % | 99,08 % | 97,24 % | 15,549 |
| ZymoBIOMICS D6300 · PromethION | 2019 · R9.4.1 | 88,77 % | 99,00 % | 97,14 % | 14,731 |
| Zymo HMW mock · PRJNA934154 current chemistry, basecalling sup |
2023 · R10.4.1 | 97,30 % | 98,53 % | 97,26 % | 28.774 |
| Degraded series from D6300 most severe level of the experiment in §6 |
2019 degraded | 74,23 % | 99,40 % | not claimed | — |
The better the chemistry, the more reads resolved to species at the same accuracy: 99 in 2015, 15,549 in 2019, 28,774 in 2023, always above 97% correct. That is the behaviour expected of an engine that respects the limit of the data: it does not overclaim when the material is poor, and it exploits the material when it is good.
Each accuracy is conditioned on the reads the system does classify, and that denominator changes between rows: 89.5% of reads in 2023 against 53.7% in 2019. That is why the reads column always travels next to the percentage, in all five rows.
Brown et al. (2017) sequenced four pure cultures and several mixtures of declared composition on MinION, and published the accuracy of MG-RAST, Kraken and One Codex on them. Paytuví-Gallart et al. (2019) reprocessed those same files with GAIA and added their column to the table. It is the only point where an OmniOta figure can be placed beside a competitor's on known ground truth: the composition of the samples is not a matter of opinion, and the competitor's figure is published.
The other third-party papers that used GAIA — dog faeces, air in Ghana, pig whey, dust in Moscow — are real samples with no known composition: they show whether two tools agree, not which one is right.
"Accuracy for GAIA was calculated as: Accuracy = # reads classified correctly / # reads classified"
Of the reads the tool dares to classify, what fraction is right. It does not penalise leaving a read unclassified; it penalises classifying it wrongly. OmniOta is measured in exactly the same way, and additionally publishes the coverage that metric does not capture.
One read per molecule is kept: the longest of those sharing a pore identifier. No true 2D consensus is built, because the repository deposits only the strands, so the work is done at a slightly worse quality than GAIA had. The reference is 22,514 genomes, one per prokaryotic species at the best assembly level available: 1,118,270 sequences, 35.1 GB compressed in nine parts, with the full lineage in each sequence header. Alignment with minimap2 -c -x map-ont -N 20 -p 0.8, one pass per part, grouping by read so that each one sees all of its alignments.
The observed identity of an alignment is (1 − ε) × evolutionary identity. Only the second carries taxonomic meaning. The sequencer error is estimated from the dataset itself — the 99th percentile of the best identity per read — and the identity, coverage and margin cut-offs are expressed as a fraction of that ceiling, never as absolute values.
The descent from domain to species is decided by consensus over the near-optimal set, with dereplication by lineage: if several labels tie, the read stays at the higher rank. No rank is claimed that is not identifiable.
A 7% variation between datasets from the same study and the same chemistry. Any absolute threshold chosen for one of them would be mis-calibrated for the rest: hence the scale is derived on every run and is not fixed in the code.
The measurement is made at genus because that is the rank Brown et al. analysed and state literally in their paper, naming exactly the datasets of this comparison. GAIA's authors took the MG-RAST, Kraken and One Codex columns from that same paper, so those are genus-level with certainty, and GAIA's own column sits in the same table compared against them. The supplementary heading says "at species level", and that ambiguity is declared here so anyone can judge it.
Nine datasets downloaded from the European Nucleotide Archive, projects PRJEB8672 and PRJEB8716. The read counts match the published ones exactly in all nine: that is the check that the starting material is the same and not a convenient subset. The four organisms in the pure cultures belong to four distinct genera from three distinct phyla.
| Dataset | Ground truth | OmniOta | Coverage | GAIA | Kraken | One Codex | MG-RAST |
|---|---|---|---|---|---|---|---|
| Ecoli | E. coli | 95,00 % | 19,8 % | 100,0 % | 99,5 % | 98,7 % | 74,7 % |
| Pfluor | P. fluorescens | 94,64 % | 40,8 % | 85,83 % | 84,6 % | 84,2 % | 84,9 % |
| Maeru | M. aeruginosa | 96,58 % | 20,6 % | 96,39 % | 85,8 % | 95,1 % | 53,1 % |
| Selong | S. elongatus | 100,00 % | 12,0 % | 100,0 % | 98,1 % | 97,6 % | 87,9 % |
| Equal (5) c | mixture of 4 | 97,70 % | 24,4 % | 93,47 % | 97,6 % | 87,4 % | 65,0 % |
| Equal (6) c | mixture of 4 | 100,00 % | 3,3 % | 98,94 % | 98,0 % | 98,7 % | 85,9 % |
| Rare (6) c | mixture of 4 | 100,00 % | 3,6 % | 100,0 % | 99,1 % | 98,7 % | 92,9 % |
| Mean (7) | 97,70 % | 17,8 % | 96,37 % | 94,67 % | 94,34 % | 77,77 % | |
| Aggregate | 896/932 reads | 96,1 % | IC 95 % [94,7 – 97,2] | — | — | — | — |
c upper bound: in a mixture, any read assigned to any of the four members counts as correct. It affects all five tools equally, and their published figures were computed on those same mixtures. GAIA's coverage is absent because their supplementary did not publish it, so that column has no counterpart and is not a comparison.
OmniOta obtains the highest mean of the five tools, and the largest advantage falls on the hardest dataset: P. fluorescens, where the four published tools land between 84.2% and 85.8% and we reach 94.64%. Both means are published: the per-dataset one because it is the only one comparable with the literature, and the aggregate one because it is the statistically correct one.
A mean of seven percentages gives the same weight to a dataset of 27 classified reads as to one of 317. It is computed that way because that is what GAIA did, and without it no comparison would be possible. The aggregate estimate is 896 correct out of 932 reads, Wilson 95% CI [94,7 % – 97,2 %], and that interval contains GAIA's 96.37%: at this data volume the mean advantage does not reach significance, and it is declared as such. The highest figure in the table is ours against all four published competitors; what would require more data is claiming that the advantage repeats on any other dataset.
The GAIA, Kraken, One Codex and MG-RAST columns are the ones published by their authors: GAIA has not been run, being a closed service with a proprietary reference database. Part of the difference between the two figures is attributable to the reference and not to the method, and there is no way to quantify it without access to theirs. The repository includes a reproduction of GAIA's rule, useful for understanding its behaviour, which is not used in this document: comparing our reproduction of their method against our method would not be a comparison between tools.
Shigella is polyphyletic within Escherichia coli — clones that acquired the invasion plasmid — and its average nucleotide identity exceeds the species threshold. The name is retained for clinical, not taxonomic, relevance. At genus rank, counting a read of E. coli assigned to Shigella as an error measures an artefact of nomenclature, not a mistake by the classifier. It affects 36 of the 78 reads in the Ecoli dataset: without the merge the figure would be 39.7% instead of 94.6%, and that is why it is declared.
The dataset staggered es HM-783D from BEI Resources, el Microbial Mock Community B of the Human Microbiome Project: twenty certified species spread between 0.03% and 41.25% of genomic DNA. It demands telling S. aureus de S. epidermidis, three Streptococcus apart from each other, and detecting members three orders of magnitude below the dominant one — which is where a classifier really breaks, as against four well-separated genera that are almost got right by elimination.
at genus, 185 of 191 classified reads, an upper bound because it is a mixture. Brown published 93% on this same material with their own pipeline. It does not enter the mean of the seven columns because GAIA published no figure for it.
Four of the twenty members have changed genus since 2010 and the current reference knows only the new name. Without declaring that equivalence the bench would have counted 20% of the mock as failure when it was a perfect hit. The ground truth is transcribed with the source's name and the equivalence lives separately: rewriting the truth with today's nomenclature would make it match our reference by construction, and the comparison would stop checking anything. A test guard pins both directions — that the renamed ones match and that congeners are not merged — because a synonym table is the convenient place to slip in exactly the merges that suit you.
ZymoBIOMICS Microbial Community Standard D6300, sequenced by Nicholls, Quick, Tang and Loman (2019) on GridION and PromethION. Ten species with composition certified in the manufacturer's datasheet: eight bacteria at 12% genomic DNA and two yeasts at 2%. 50,000 reads measured per platform, from a prefix yielding more than 200,000 — against the 5,600 2D reads of the whole of Brown 2017. It contributes two axes the first material could not give: manufacturer certification and a measurement of invariance between sequencers.
| GridION | PromethION | |
|---|---|---|
| reads assigned to genus | 29,824 | 27,928 |
| accuracy at genus | 99,08 % | 99,00 % |
| reads assigned to species | 15,549 | 14,731 |
| accuracy at species | 97,24 % | 97,14 % |
| largest false positive | 0,17 % | 0,21 % |
It is a mixed community, so accuracy is an upper bound: any read assigned to any member counts as correct. The reads come from the run's prefix, which over-represents the flow cell's first hours; for comparing between platforms the bias acts the same way in both, and for an absolute figure it comes from the best stretch of the run and is stated as such. The two yeasts come out at exactly 0.00% on both platforms: the reference was built from RefSeq's bacterial catalogue and contains no fungi, and the figure being identical on both confirms it.
The same community, the same preparation, two different sequencers. The question is answered separately for the certified community and for everything emitted, because the answers differ and giving only one of the two would hide something.
| Community | Everything emitted | |
|---|---|---|
| observed L1 | 0,0181 | 0,0231 |
| L1 expected from sampling (p95) | 0,0257 | 0,0276 |
| divergent taxa | 0 | 1 |
| verdict | Invariant | no |
The composition is invariant between GridION and PromethION. The per-genus differences run from 0.05 to 0.43 points and the global divergence is practically the one sampling itself produces: not "close to zero", but statistically indistinguishable from chance, which is what had to be shown.
What does not repeat are the low-abundance artefacts: Macrococcus, which is not a member of the community, appears with 23 reads on one platform and 5 on the other. The composition repeats; the traces do not. It is a direct operational warning about how much weight to give a taxon held up by twenty reads, and it agrees with the Wilson doctrine the product already applies to per-taxon confidence.
The D6300 datasheet declares five composition columns for the same community: genomic DNA, 16S, 16S&18S, genome copies and cells. It is even in DNA — 12% per bacterium — and deeply uneven in 16S: P. aeruginosa goes from 12% to 4,2 % and L. fermentum del 12 % al 18,4 %, a factor of ×4.4 between the two. And S. cerevisiae is 9.3% of the 16S&18S and 0.29% of the cells: two correct answers to two different questions. Comparing a shotgun result against the 16S column would produce a bias of that magnitude chargeable to the bench and not to the engine — it is exactly what genome-size normalisation corrects in production, and here is the material that proves it beyond argument.
A 2026 product evaluated on 2019 chemistry is a legitimate objection, and it is not answered with arguments but by measuring. ZymoBIOMICS HMW DNA Standard, ONT R10.4.1 at 5 kHz with Kit 14, study PRJNA934154 from 2023. Identity is measured by aligning against the standard's reference genomes; it is not taken from the manufacturer's specification sheet.
| Material | Chemistry | Measured identity |
|---|---|---|
| Brown et al. · 2015 | R7.3 / R9 | 74,23 % |
| Nicholls et al. · 2019 | R9.4.1 | 88,77 % |
| Zymo HMW mock · 2023 | R10.4.1 (sup) | 97,30 % |
Between the material previously used and the current chemistry there are 8.5 points of identity. That is not a nuance: it is the difference between being near the wall of an exact k-mer method and being far from it. Percentiles of the modern material: p25 = 95.14%, p75 = 98.16%, p95 = 98.92%, p99 = 99.33%.
| Genus | Species | |
|---|---|---|
| classified reads | 89,5 % | 57,5 % |
| accuracy | 98,53 % | 97,26 % |
| standard members found | 7/7 | 7/7 |
Against the 2019 material, reads resolved to species go from 31.1% to 57,5 % at the same accuracy (97.24 → 97.26%). On current chemistry the system resolves nearly twice as many reads to species without losing accuracy. Genus accuracy is half a point below the 2019 figure for the inverse reason: here 89.5% of the reads enter instead of 53.7%, so the hard ones enter too.
The same run is published with all three ONT basecallers. Same raw signals, same pore, same sample, same day: the only variable is the software that translates signal into bases.
| Basecaller | Identity | Aligned |
|---|---|---|
| fast | 89,80 % | 9.423 |
| hac | 96,16 % | 9.456 |
| sup | 97,30 % | 9.464 |
A 2023 run with fast basecalling lands at the same identity as the 2019 material with the best basecalling.
Low identity is not a historical condition the industry has left behind: it is a state entered today, through an ordinary operational decision — using the fast basecaller because sup costs hours of GPU — and also with DNA degraded by thermal processing, a fatty matrix that is hard to extract from, or a flow cell at the end of its life. fast consumes 7.5 of the ~15 points of margin separating the modern material from the point where an exact k-mer method stops returning a result. Half the cushion, for one checkbox.
On the 16S/amplicon axis the bench already covers R10.4.1 chemistry with material from September 2025 (MSPlus community, ont-open-data/zymo_16s_2025.09): 12 of 12 members detected, zero regulated false positives, one over-claim. Measured and recorded in the taxonomic-resolution documentation, and cited here as what it is: a measurement from another session, not re-run in this one.
Between Brown 2017 and Zymo D6300 everything changes at once: the year, the chemistry, the sequencer, the community, the laboratory and the identity of the reads. With two points differing in six variables, attributing a collapse to one of them is not justified, and any reviewer's obvious retort would be the right one: you picked an old dataset.
This experiment exists so that sentence stops being arguable. It takes the same file from Zymo and injects error at increasing rates, with a fixed seed and three depths to separate identity from depth. Everything else is identical by construction: the only free variable is identity. And it could have gone against us: had the collapse failed to reproduce, the published explanation would be false and would have to be withdrawn. That is why it counts.
Substitutions only, no insertions or deletions. Real nanopore error has indels, and they are más destructive for an exact k-mer, because they shift the frame of the 31 bases instead of breaking one position. Omitting indels makes the experiment conservative in the competitor's favour: it is given the more benign of the two scenarios. Identity is measured by aligning, not assumed: on material at 88.77%, injecting 20% does not give 80% but 74.23%.
The most degraded level lands at 74,23 %, which is the median identity of Brown 2017. The experiment reaches by construction the very point it set out to explain, starting from modern, certified material: no 2015 dataset was needed; degrading the 2019 one is enough.
| Identity | Members OmniOta |
Accuracy at genus |
Taxa sylph |
Genera spurious |
|---|---|---|---|---|
| 88,77 % | 8/8 | 99,07 % | 8 | 0,93 % |
| 87,17 % | 8/8 | 99,16 % | 7 | 0,84 % |
| 84,91 % | 8/8 | 99,29 % | 6 | 0,71 % |
| 82,65 % | 8/8 | 99,18 % | 0 | 0,82 % |
| 79,68 % | 8/8 | 99,44 % | 0 | 0,56 % |
| 76,84 % | 8/8 | 99,37 % | 0 | 0,63 % |
| 74,23 % | 8/8 | 99,40 % | 0 | 0,60 % |
sylph column at 5,000 reads; at 20,000 and 50,000 the collapse shifts but reproduces just the same. The ground truth is 8 members at all seven levels.
All eight organisms of the community are recovered at every level, including the most degraded — the same one where the competitor does not return a single line. It is the result that matters for a real food sample: DNA fragmented by thermal processing, a fatty matrix, difficult extraction.
Accuracy rises as the material degrades, and that is not a merit: it is selection bias. It is conditioned on the reads the system does classify; as the material degrades the doubtful ones stop clearing the threshold and the easy ones remain, so the hit rate over the survivors improves by selection, not by capability. The figure that holds this section up is members found — 8/8 — together with classified reads, which fall from 53.7% to 36.6%.
Against GAIA only its published figures are available: it is a closed service with a proprietary reference database, so not even running it would remove the reference asymmetry. With sylph (Shaw & Yu, Nature Biotechnology 2024) it can be done: it was given exactly the same 22,514 genomes, one by one, with the index built in 296 s. And MetaPhlAn 4.2.6 was run on the same file with its own marker database, the only one it accepts.
| OmniOta | sylph | |
|---|---|---|
| per-read accuracy at genus | 99,18 % | not applicable |
| community members found | 8/8 | 8/8 |
| taxa emitted | 61 | 8 |
| weight of the surplus taxa | 0,82 % | — |
| profiling time | ~75 min | 24 s |
Both tools find the complete community, and each is stronger on a different axis. Per read OmniOta is highly precise: 99.18%, and that is the figure governing an abundance table and everything computed on top of it. Per named taxon, sylph decides better: it emits 8 where we emit 61, because it models genome coverage with k-mer statistics and requires the genome to be genuinely present before naming it, whereas we name any taxon that receives a read assignment.
The 53 surplus taxa add up to 0.82% of the reads —Microbacterium 51, Streptomyces 43, Macrococcus 23, and a tail below ten reads each — and the criterion under development is a minimum evidence to name a taxon, not merely to assign reads to it: the same Wilson doctrine the product already applies to per-taxon confidence, carried into the decision to publish it. It is consistent with what §4.2 measures: composition repeats between platforms, traces do not.
The two times are not directly comparable: the ~75 min include aligning against the nine parts of the reference while building the index on every pass, whereas sylph reuses an index built once. With a persistent index the difference would shrink considerably, and that measurement is pending. sylph is a profiler: it gives the list of genomes present with their abundance and does not assign each read, which is why it has no per-read accuracy column. Shigella boydii in its output is not a false positive: it is the Escherichia/Shigellacomplex, counted as correct by the same criterion applied to OmniOta in §3.1.
Kraken2/Bracken is not measured: its database with these same 22,514 genomes would require on the order of 170 GB of RAM, as measured, so the reason is published instead of giving it a trimmed database that would distort the comparison.
The certified composition of the HMW standard is equimolar between bacteria: 14.3% of genomic DNA each. Counting reads, Salmonella gathered 47.1% and Bacillus 3.2%. It is not subset bias — the profiles of the first and last 10,000 reads are practically identical — it is the metric. Counting reads is not counting DNA mass. Abundance is mean coverage, aligned bases divided by genome size, which is what CoverM, sylph and CAMI hold to.
| Organism | By reads | Coverage | MetaPhlAn | Cert. |
|---|---|---|---|---|
| S. enterica | 47,1 % | 25,2 % | 25,6 % | 14,3 % |
| E. faecalis | 11,9 % | 21,3 % | 23,3 % | 14,3 % |
| S. aureus | 4,8 % | 16,9 % | 18,9 % | 14,3 % |
| L. monocytogenes | 4,0 % | 13,0 % | 9,3 % | 14,3 % |
| E. coli | 24,0 % | 10,7 % | 13,7 % | 14,3 % |
| B. subtilis | 3,2 % | 7,6 % | 4,2 % | 14,3 % |
| P. aeruginosa | 4,9 % | 5,3 % | 4,9 % | 14,3 % |
| L1 divergence | 0,852 | 0,411 | 0,501 | — |
Measuring DNA mass instead of counting reads halves the error — 51.8% less — on material with certified composition. Counting reads, Salmonella appeared to be 3.3 times its real mass and Bacillus a quarter of its own. The rule is not a matter of style.
The comparison with MetaPhlAn counts precisely because the two tools share nothing: we align against 22,514 complete genomes, MetaPhlAn profiles against its clade-specific UniRef markers. The two methods agree to within 2.14 points on average and on the ordering of the three largest, so the sequenced material is not equimolar, and that is established by two independent routes. And on what it does compete on, OmniOta lands 18% closer to the certified composition: L1 0.411 against 0.501.
For the remaining divergence there are at least two known reference-side causes before invoking any from the method: the "E. coli" in the panel is the genome of Shigella boydii and measuring coverage against a surrogate genome underestimates it; the same applies to P. aeruginosa PAO1. There is room for real extraction and lysis bias, which this experiment does not separate from the above. MetaPhlAn detects S. cerevisiae and we do not, because our reference is bacterial: it is a capability of theirs that we lack.
| document version | 1.4.0 |
| date of measurement | 2026-08-10 |
| engine | minimap2-lca/3.0.0 |
| working tree | 61a0bb7c |
| reference | 22,514 genomes · 1,118,270 sequences · 35.1 GB |
| aligner | minimap2 · map-ont · N 20 · p 0.8 |
| competitors run here | sylph · MetaPhlAn 4.2.6 |
The procedures for downloading the material, building the reference and measuring are published in the repository as executable commands, and every methodological decision in this document is pinned by an automatic test guard that compares the text against the run's JSON: if anyone changes a figure in the prose without measuring again, the test fails.
The full detail of each axis — extended caveats, per-dataset results, the bench's version history and the figures withdrawn for not being reproducible — lives in the repository's canonical documentation, available under agreement for technical evaluation.