DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome
-
Updated
Jan 22, 2026 - Python
DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome
Quickly search, compare, and analyze genomic and metagenomic data sets.
Inference of ploidy and heterozygosity structure using whole genome sequencing data
Accurate metagenomic profiling && Fast large-scale sequence/genome searching
De novo genome assembly and multisample variant calling
modular k-mer count matrix and Bloom filter construction for large read collections
A versatile toolkit for k-mers with taxonomic information
Generate unique KMERs for every contig in a FASTA file
Fast k-mer based tool for multi locus sequence typing (MLST)
Ultra-rapid detection of viral variants directly from sequencing data
Fast and space-efficient taxonomic classification of long reads
Predict plasmids from uncorrected long read data
Bioinformatics 101 tool for counting unique k-length substrings in DNA
Generate kmers/minimizers/hashes/MinHash signatures, including with multiple kmer sizes.
k-mer similarity analysis pipeline
Count kmers with a more efficient (faster) hash table
Topsicle utilizes abundance of telomere pattern k-mers to estimate telomere length in long read.
To associate your repository with the kmer topic, visit your repo's landing page and select "manage topics."