HelixCore · Precision Genomics. Unlimited Power. The twelve modules
FILE BM–2026–08
ENGINE integrated pipeline v3
CLASS. TECHNICAL · PUBLIC
METHODS 6 orthogonal
GATES veto by threshold
06 Discovery · diagnostic targets

BioMiner

It finds the region of the genome that tells your organism apart from everything else, crossing six methods that share no assumptions. What shows up by two independent routes is not a coincidence: it is a target.

This is the work that takes an expert bioinformatician weeks of chained analyses — and of deciding by eye which result to believe — resolved into a ranked list with the evidence for each candidate in plain sight.

How to read this page
Every method declares its confidence level and every candidate carries which methods it came from. The score is not a loose number: it breaks down.
6 orthogonal methods · cross consensus
9 steps in a single integrated pipeline
6
discovery methods sharing no assumptions, crossed by automatic consensus
2+
concurring methods trigger an exponential confidence boost on the region
9
steps in a single pipeline: they used to be separate pipes you had to cross by hand
0–1,0
inclusivity: the fraction of target genomes where the region is actually present
What makes it unique
The k-mer analyser performs peak calling over k-mer space, aligning nothing. That captures diagnostic regions in rearranged genomes which aligners miss entirely — and it is a line of evidence no commercial suite ships.
On top of that, consensus: a region found by two or more methods that share no assumptions receives an exponential confidence boost, because agreeing by accident in two different spaces — alignment, variants, k-mers, pangenome — is far less likely than getting one right.
01 The state of the art

One method only says what that method can see.

Every approach has a known blind spot: the aligner misses rearrangements, the pangenome does not see the point variant, the variant caller does not see what is missing entirely. Choosing one is choosing its blind spot.
Tool
How far it goes
What BioMiner adds
Geneious, CLC Genomics Workbench
Alignment, annotation and genome comparison in a comfortable graphical environment.
Adds what is done by eye there: k-mers with peak detection, pangenome and automatic consensus across six orthogonal methods, with the evidence cross-check scored by the system.
Ridom SeqSphere and typing suites
They compare isolates over defined, well-validated allele schemes.
Needs no published scheme: it discovers candidate regions on your own genomes, including those no scheme covers yet.
Roary, Panaroo, Snippy on their own
Each answers its own question very well: pangenome, variants, structure.
Runs them under common coordinates and a single scoring criterion, so weeks of integration become a reproducible consensus.
Picking the target from the literature
It is quick and uses a gene that already worked for another laboratory.
Checks inclusivity against your collection: that the target is present in all your strains and absent from your exclusions, before it reaches the bench.
02 The six methods

Genuinely orthogonal: each sees what the others do not.

And each enters the ranking with its own declared confidence: variant evidence weighs more than structural, and structural more than k-meric. Not all routes are worth the same, and that is written down.
01 Structural
What is in the target that is not in the rest
Multi-genome alignment to detect exclusive and shared regions, with automatic rerouting to similarity search when the sequence is short.
minimap2 + BEDTools · confianza 0,8
02 Variants
SNPs and clustered insertions and deletions
Dual-route detection depending on the material: one for simple genomes and another for assemblies in many contigs. Nearby variants are grouped into diagnostic windows.
Snippy · minimap2 + bcftools · confianza 0,9
03 K-mers
Peak calling without aligning anything
An alignment-free method with peak calling over k-mer space. It captures diagnostic regions in rearranged genomes, where the aligner simply finds nothing to align.
jellyfish + MinHash + UPGMA · confianza 0,7
04 Pangenome
Which genes belong to everyone and which are only yours
Gene prediction and classification by population prevalence into core, shell and cloud, to identify genes exclusive to the target and absent from the exclusions.
Prodigal + CD-HIT · Roary o Panaroo
05 Inclusivity
The target is in ALL your genomes, not in one strain
Every candidate is checked against the full set of the target organism. This is the check that prevents picking a region specific to a single isolate and finding out in the laboratory.
Score 0–1.0 by fraction of genomes
06 Ranking
Consensus, designability, robustness and confidence
A composite score over four dimensions plus an optional primer-quality one, with the exponential boost for anything found by several methods. Every component is visible separately.
Optional micro-GWAS: Fisher and Mann-Whitney
03 How it works

A single pipeline, nine steps, and two gates at the end.

The integrated pipeline replaced the isolated ones: you used to launch the structural analysis on one side and the variants on the other, and cross the results by hand.
01
Input and enrichment
Your own files or public accession numbers, which the system downloads and enriches with their metadata: GC content, length, organism and source annotation.
02
Quality control and structure
Read cleaning and multi-genome alignment against the declared exclusions, which is where the map of the exclusive and the shared is drawn.
03
Variants, k-mers and pangenome
All three analyses run over the same material and in the same coordinate system, which is what makes it possible to cross them afterwards without translating formats.
04
Discrimination engine
It ranks candidates by inter-method consensus, designability, robustness against multiple exclusions and the confidence of the method that found them.
05
Quality gates and export
Thresholds with veto power mark each candidate as passed or rejected, and the set leaves enriched — sequence, metadata and scores — towards primer design.
A counter that does not lie
The library keeps three populations in the same place: the material you uploaded, the contigs an assembly breaks into, and the artefacts the analysis manufactures. A dataset with two assemblies, their 65 contigs and 3,054 pangenome genes was announcing 3,121 sequences. The figure that means something — how many organisms are inside — is 2.
Fixed with a single rule in the code and its guardian test. A dataset with no material counts zero, not the number of artefacts: that is the correct answer and it also exposes orphaned datasets.
The quality gates
Every candidate passes thresholds with veto power: a region may have three-method consensus and still not be designable. The gate stops it there, and says why, instead of letting it reach primer design and fail later and more expensively.
04 The contract

What goes in, what comes out, what it chains to.

In
Genomes of the target organism and of those to be excluded: your own files, public accession numbers, HoloGen assemblies, or taxa exported from OmniOta with their real sequence.
Out
A ranked list of candidate regions with the score broken down, the methods that found each one, its inclusivity, the quality-gate verdict and the sequence ready to design on.
Chains to
Direct export to SnapPrime for primer design and to FastKit for the kit, with the candidate's traceability intact: which genomes it came from and by which methods.
05 Technical sheet
Input
Your own FASTA, accession numbers from public databases, HoloGen assemblies, or taxa exported from OmniOta with their real sequence. Target genomes and exclusion genomes, declared separately.
Output
A ranked set with the score broken down per candidate, a genome map with one track per method, a consensus report and export with full traceability.
Confidence by method
Variant evidence weighs 0.9; structural 0.8; k-meric 0.7. A consensus between methods of different weight is not averaged blindly.
Inclusivity
The fraction of target genomes in which the region is present, from 0 to 1. A candidate with high consensus and low inclusivity is flagged: it is strain-specific, not organism-specific.
Material counter
The library counts organisms, not rows: the contigs of an assembly and the artefacts of the analysis are not dataset entries. A single rule in the code, with a guardian test.
Limit · ranking
The score is heuristic and is not yet expressed as curated-knowledge rules: the breakdown by component is visible, but not a citable rule behind each weight. It sits in the module's debt register.
Limit · retry
The integrated pipeline is all or nothing: there is no per-step retry nor intermediate checkpoint, so a late failure forces the whole analysis to be repeated.
Limit · pangenome
The pangenome is optional and switching it on is left to the user's judgement: with few or very close genomes it adds little, and the interface does not yet guide that decision.
Structural alignment: minimap2 and BEDTools, with automatic rerouting to BLAST for sequences under 5 kb. Variants: Snippy for simple genomes, minimap2 and bcftools for multi-contig. K-mers: jellyfish, with MinHash distance and UPGMA clustering. Pangenome: gene prediction with Prodigal and clustering with CD-HIT, Roary or Panaroo. Optional micro-GWAS: Fisher's exact test and Mann-Whitney with Bonferroni correction and false discovery rate.

Six methods, one consensus: the most robust target, not the easiest one.

Request access See the twelve modules
BIOTECNO.org · Vitoria-Gasteiz EU + Codex
© 2026 BIOTECNO Research Group
Legal noticePrivacyCookies