FlavoTyper Troubleshooting¶
This document covers the most common errors and questions encountered when installing or running FlavoTyper.
Table of contents¶
- Getting started
- First time with conda?
- First time with PyPI?
- Installation issues
ERROR: blastn not found in PATH/ERROR: makeblastdb not found in PATHERROR: fastANI not found in PATHmakeblastdb failed- Input genome errors
- No FASTA records parsed from file
- Malformed FASTA — duplicate contig names
- Invalid nucleotide characters in genome
- QC and species check
- Sample is reported as
NotTyped— species check failed - fastANI subprocess failure
- Genome size warning in
QC_warnings - High contig count warning in
QC_warnings - Typing results
Call_stateisPartial— one component isUndefinedCall_stateisUndefined— nothing was typedCall_stateisAmbiguous— multiple interpretations matched- Warning: markers detected on different contigs
- Warning: R typing remained undefined — R-associated marker detected
- Warning: novel serotype — no reference strain available
- Locus analysis
- Locus map not generated despite
--locus-analysis - Warning: locus comparison failed for sample
- Other
ERROR: Duplicate sample names detectedERROR: Unsupported genome extensionERROR: Locus reference database not found- Custom database errors (
--db)
Getting started¶
New to conda or Python virtual environments? These walkthroughs take you from a fresh laptop to a working FlavoTyper install. They expand on the commands in the Installation section of the README.
First time with conda?¶
Bioconda packages are installed with conda. If you don't already have a conda-based package manager, install one first (step 1); otherwise skip to step 2.
Step 1 — Install a conda package manager
We recommend Miniforge, a minimal installer that comes with both conda and
the faster mamba, and is pre-configured for the conda-forge channel.
Linux / macOS:
# Download the installer for your OS and CPU architecture
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
# Run it and follow the prompts (accept the license, accept the default location)
bash "Miniforge3-$(uname)-$(uname -m).sh"
# Close and reopen your terminal so the installation takes effect
Verify it worked:
conda --version
Lightweight alternative — micromamba: if you prefer a single self-contained binary with no base environment, run
"${SHELL}" <(curl -L micro.mamba.pm/install.sh)and reopen your terminal. In the commands below, replacecondawithmicromamba.
Step 2 — Create and activate an environment
This makes an isolated environment named flavotyper so nothing conflicts with
your system, and installs FlavoTyper into it:
conda create -n flavotyper -c conda-forge -c bioconda flavotyper
conda activate flavotyper
This single command installs FlavoTyper and its external tools (BLAST+,
fastANI) — no further steps needed. (Tip: swap conda for mamba for a faster
install.)
You'll need to run conda activate flavotyper once in each new terminal session
before using the tool. Then confirm everything works:
flavotyper --version
flavotyper data-dir
First time with PyPI?¶
PyPI installs the FlavoTyper Python package only; you install the external tools (BLAST+, fastANI) separately afterwards.
Step 1 — Check you have Python 3.10 or newer
python3 --version
If the command is missing or the version is below 3.10, install Python from
python.org/downloads (or, on Debian/Ubuntu:
sudo apt install python3 python3-venv python3-pip).
Step 2 — Create and activate a virtual environment
A virtual environment (venv) keeps FlavoTyper's Python dependencies separate
from your system Python:
python3 -m venv .venv # create it (once)
source .venv/bin/activate # activate it — Windows: .venv\Scripts\activate
Your prompt will now show (.venv). You'll need to re-run the activate command
in each new terminal session.
Step 3 — Install FlavoTyper
python3 -m pip install --upgrade pip # ensure pip is current
pip install flavotyper
Step 4 — Install the external dependencies
PyPI does not install BLAST+ or fastANI. The simplest way to get them is conda:
conda install -c conda-forge -c bioconda blast fastani
Then confirm everything works:
flavotyper --version
blastn -version
fastANI --version
Installation issues¶
ERROR: blastn not found in PATH / ERROR: makeblastdb not found in PATH¶
BLAST+ is not installed or not on your PATH. Install it using one of the following methods:
conda / mamba
conda install -c bioconda blast
apt (Debian / Ubuntu)
sudo apt-get install ncbi-blast+
macOS (Homebrew)
brew install blast
Manual download
Pre-compiled binaries for Linux, macOS, and Windows are available from the NCBI BLAST+ download page. Download the archive for your platform, extract it, and add the bin/ directory to your PATH.
Verify the installation:
blastn -version
makeblastdb -version
ERROR: fastANI not found in PATH¶
fastANI is not installed or not on your PATH. Install it using one of the following methods:
conda / mamba
conda install -c bioconda fastani
apt (Debian / Ubuntu)
sudo apt-get install fastani
macOS (Homebrew)
brew install fastani
Manual download Pre-compiled binaries are available from the FastANI GitHub releases page. Download the binary for your platform and place it somewhere on your PATH.
Verify the installation:
fastANI --version
If you cannot install fastANI, disable the species check:
flavotyper type --genomes genome.fasta --outdir results/ --no-species-check
makeblastdb failed¶
BLAST+ is installed but the internal marker database could not be built at runtime. The full error from BLAST is included in the message. Common causes:
- Insufficient disk space in the system temporary directory.
- Permissions issue preventing writing to the temp directory.
Free up disk space or set a writable temp directory, then retry.
Input genome errors¶
No FASTA records parsed from file¶
The input genome file is empty, unreadable, or not a valid FASTA file. Verify that the file is not empty and starts with a > header line.
Malformed FASTA — duplicate contig names¶
Two or more contigs in the same genome file share the same identifier. This is a malformed assembly. Re-run the assembler or rename the duplicate contigs before resubmitting.
Invalid nucleotide characters in genome¶
The genome sequence contains characters outside the allowed IUPAC DNA codes (A/C/G/T/N/R/Y/S/W/K/M/B/D/H/V). This can happen with protein FASTA files submitted by mistake, or assemblies exported with non-standard placeholders. Verify the file is a nucleotide assembly.
QC and species check¶
Sample is reported as NotTyped — species check failed¶
The input genome did not reach the ANI threshold (default: 95 %) against the F. psychrophilum type strain (NCIMB 1947T) used as a reference. This usually means:
- The genome does not belong to F. psychrophilum — verify species identity independently.
- The assembly is very fragmented or contaminated — inspect
N_contigsandGenome_size_bpin the output.
If species identity has been confirmed by other means, bypass the gate:
flavotyper type --genomes genome.fasta --outdir results/ --no-species-check
fastANI subprocess failure¶
fastANI is installed and found, but crashed during execution. The error message includes the fastANI stderr output. Common causes:
- fastANI version incompatibility — the minimum required version is 1.3.
- The genome file is valid FASTA but too short or too fragmented for fastANI to process.
- No ANI hits were produced against the reference — this results in a failed species check with the message
No fastANI hits against provided species references.
Verify the fastANI version (fastANI --version) and inspect the genome file. If the genome is genuinely F. psychrophilum and the failure is a fastANI compatibility issue, use --no-species-check.
Genome size warning in QC_warnings¶
The genome size falls outside the expected interval [2,619,202 – 3,122,663 bp]. This is an advisory warning — typing still proceeds. Possible causes:
- Incomplete assembly (too small) or contamination with non-target sequences (too large).
- The genome genuinely falls outside the reference interval — this is not uncommon for divergent strains.
The thresholds can be adjusted:
flavotyper type --genomes genome.fasta --outdir results/ \
--expected-genome-size-min 2500000 \
--expected-genome-size-max 3300000
High contig count warning in QC_warnings¶
The assembly has more contigs than the advisory threshold (default: 300). Typing still proceeds, but highly fragmented assemblies may split marker genes across contigs, which can affect distance-rule evaluation for R1V2. Improving assembly contiguity is recommended before relying on variant-level calls.
Typing results¶
Call_state is Partial — one component is Undefined¶
Either the O-type or the R-type could not be assigned. Common causes:
- O-type Undefined: no wzy marker was detected above the identity and coverage thresholds. Check
Present_markers— if the wzy hit is present inRaw_hitsbut absent inPresent_markers, it likely fell just below the threshold. Consider lowering--min-identityor--min-coveragewith caution. - R-type Undefined: no R base-group marker was detected. This is expected for O:0 (which defaults to R0). For other O-types it may indicate a novel R type or a fragmented assembly missing the R locus.
Check Typing_warnings for additional context.
Call_state is Undefined — nothing was typed¶
Both the O-type and the R-type are Undefined and no S1 markers were detected (so S-type defaulted to S0), giving the serotype Undefined-S0-Undefined. Unlike a Partial call, no serotype component was positively assigned at all. This is distinct from NotTyped, which means the genome failed QC and never entered typing. Common causes:
- A novel or divergent strain whose O-antigen and R-region markers are not represented in the database.
- A poor or fragmented assembly missing the LPS biosynthesis locus.
- An off-target genome that passed the ANI species gate but is not a typeable F. psychrophilum.
Inspect Raw_hits to see whether any markers were detected near (but below) the thresholds, and check assembly quality (N_contigs, Genome_size_bp). Lowering --min-identity / --min-coverage may help if hits fell just short, but interpret such calls with caution.
Call_state is Ambiguous — multiple interpretations matched¶
More than one O-type marker or more than one R base group was detected. Possible causes:
- Contamination with additional F. psychrophilum strains carrying a different serotype.
- Assembly artefacts (chimeric contigs, misassembly near the O-antigen locus).
- Highly fragmented assembly causing spurious repeat hits.
Inspect Present_markers, Alternative_serotypes, and Typing_warnings.
Warning: markers detected on different contigs¶
The typing markers are split across multiple contigs. Typing still proceeds, but the distance rule for R1V2 (requiring Rieske and wfpF_p to be within −6 to +6 bp) cannot be reliably evaluated. A Resolved call is still possible if marker presence alone is sufficient, but the variant assignment may be less certain.
Warning: R typing remained undefined — R-associated marker detected¶
An R-associated marker was detected but did not satisfy any complete base-group rule. This can happen when:
- The strain carries a partial R locus (e.g. only one of the required markers is above the threshold).
- The strain has a novel R type whose markers are not yet fully characterized in the database.
Check Present_markers and Raw_hits to see which marker was detected and at what identity/coverage.
Warning: novel serotype — no reference strain available¶
The resolved serotype (e.g. O:3-S1-R4) is not present in the built-in reference locus database. Typing is still reported, but no locus analysis can be performed. This is expected for newly described serotype combinations. Please consider reporting the strain to the FlavoTyper maintainers so the database can be extended.
Locus analysis¶
Locus map not generated despite --locus-analysis¶
Locus analysis only runs when all three conditions are met:
--locus-analysisis enabled.- The call state is
Resolved. - The resolved serotype has a reference entry in the built-in locus database.
If the call is Partial or Ambiguous, or if the serotype is novel (see warning above), no locus analysis is performed.
Warning: locus comparison failed for sample¶
This warning is printed to stderr during the run:
WARNING: locus comparison failed for <sample>: <error details>
The typing call is still written to the output files — only the locus analysis artifacts (PNG, alignment, FASTA) are missing for that sample. Causes can include disk space exhaustion, a corrupted locus reference FASTA, or an unexpected BLAST failure. Check the full traceback printed below the warning for details.
Other¶
ERROR: Duplicate sample names detected¶
Two or more input files share the same filename stem (e.g. run_a/sample.fasta and run_b/sample.fasta). FlavoTyper rejects this by default to avoid output collisions.
Either rename the files to have unique stems, or allow duplicates (FlavoTyper will append a numeric suffix to the sample key):
flavotyper type --genomes *.fasta --outdir results/ --allow-duplicate-sample-names
ERROR: Unsupported genome extension¶
Input files must have a recognised FASTA extension: .fa, .fna, .fasta, .fas (optionally .gz compressed). Rename the file or verify the extension.
ERROR: Locus reference database not found¶
The built-in locus FASTA could not be found. This should not happen after a clean install. Try:
flavotyper data-dir
to confirm the data directory is present. If using a custom --locus-db path, verify the file exists at that location.
Custom database errors (--db)¶
When using a custom typing rules database via --db, FlavoTyper validates the YAML file and the marker FASTA it references before running. Errors at this stage indicate a problem with the custom database file, not with the input genomes. Common messages:
Database file not found— the path passed to--dbdoes not exist.Marker FASTA not found— themarker_fastapath declared inside the YAML does not exist relative to the YAML file.Database missing required key: <key>— the YAML is missing a required section.Unknown marker '<name>' referenced in <context>— a rule references a marker name not present in the marker FASTA.
Refer to the built-in Flavotyper_rules.yaml for the expected database structure.