A technical reference for the major file formats used in genomics, sequence analysis, and bioinformatics pipelines — with format specifications, examples, and usage guidance.
Bioinformatics relies on standardized file formats that allow genomic data to be processed consistently across different software tools, pipelines, and institutions. Understanding these formats is essential for working with sequencing data, variant calls, genome annotations, and alignment results. This reference covers the six most important formats you will encounter in modern genomics workflows, with annotated examples for each.
Sequence storage — DNA, RNA, and protein sequences
FASTA is the simplest and most universally used format for storing biological sequences. Named for the FASTA sequence alignment program (1985), it stores one or more sequences with associated metadata in a plain text format that is both human-readable and machine-parsable.
Each sequence entry consists of:
> (greater-than symbol), followed by a sequence identifier and optional description| Code | Bases Represented | Meaning |
|---|---|---|
A | A | Adenine |
C | C | Cytosine |
G | G | Guanine |
T | T | Thymine (DNA) / Uracil (RNA) |
N | A, C, G, T | Any nucleotide (unknown) |
R | A, G | Purine |
Y | C, T | Pyrimidine |
W | A, T | Weak |
S | G, C | Strong |
Raw sequencing reads with per-base quality scores
FASTQ extends FASTA by including per-base quality scores alongside the sequence. It is the standard output format of next-generation sequencing (NGS) instruments — Illumina, Oxford Nanopore, PacBio platforms all produce FASTQ output. Every short read sequencing experiment begins with FASTQ files.
Each read is represented by exactly 4 lines:
@, followed by sequence identifier and optional description (instrument, flowcell, coordinates)+ (optionally followed by the identifier again)Quality scores (Q-scores) quantify the probability of a base call error. The relationship is: Q = -10 × log₁₀(P), where P is the error probability.
| Q-Score | Error Probability | Accuracy | ASCII Character (Phred+33) |
|---|---|---|---|
| 10 | 1 in 10 | 90% | + (ASCII 43) |
| 20 | 1 in 100 | 99% | 5 (ASCII 53) |
| 30 | 1 in 1,000 | 99.9% | ? (ASCII 63) |
| 40 | 1 in 10,000 | 99.99% | I (ASCII 73) |
| 40+ (Q41) | <1 in 10,000 | >99.99% | J (ASCII 74) |
Modern Illumina sequencing typically achieves Q30+ for 75–90% of bases. Reads are typically quality-filtered (remove low-quality bases/reads) and adapter-trimmed before downstream analysis using tools like Trimmomatic, fastp, or Cutadapt.
Most Illumina sequencing is paired-end: each fragment is sequenced from both ends, producing two FASTQ files (R1 and R2) where reads at the same position in each file are paired. Paired reads are more informative for alignment and structural variant detection.
Genetic variants: SNPs, indels, structural variants
VCF (Variant Call Format) is the standard format for storing genetic variants identified from sequencing data — single nucleotide polymorphisms (SNPs), insertions/deletions (indels), and structural variants. It is produced by variant calling tools such as GATK HaplotypeCaller, DeepVariant, and Strelka2, and used by major variant databases (dbSNP, ClinVar, gnomAD).
| Column | Description |
|---|---|
CHROM | Chromosome (chr1, chrX, chr17...) |
POS | 1-based position of the first base of the variant |
ID | dbSNP rs identifier, or "." if unknown |
REF | Reference allele at this position |
ALT | Alternate allele(s) — comma-separated for multi-allelic sites |
QUAL | Phred-scaled quality score for ALT allele call |
FILTER | PASS if variant passes all filters; filter name if failed |
INFO | Semicolon-separated key=value annotations (depth, AF, functional impact) |
FORMAT | Colon-separated list of per-sample format fields (GT, GQ, DP, AD) |
SAMPLE | One column per sample with colon-separated values matching FORMAT |
0/0 — Homozygous reference (no variant)0/1 — Heterozygous (one reference, one alternate allele)1/1 — Homozygous alternate0|1 — Phased heterozygous (paternal|maternal allele)./. — Missing data (insufficient coverage)Genomic intervals and regions — gene locations, peaks, regulatory elements
BED (Browser Extensible Data) format is the standard way to represent genomic intervals — specific chromosomal regions defined by start and end coordinates. It is used for gene annotations, ChIP-seq peaks, ATAC-seq open chromatin regions, CpG islands, repeat elements, regulatory regions, and virtually any genomic feature defined by position.
| Field | Required | Description |
|---|---|---|
chrom | Yes | Chromosome name (chr1, chrX) |
chromStart | Yes | Start position, 0-based |
chromEnd | Yes | End position, exclusive (0-based) |
name | No | Feature name/identifier |
score | No | Score 0–1000 (peak signal, conservation score) |
strand | No | + (forward), - (reverse), or . (unknown) |
Sequence alignment/map — aligned reads against a reference genome
SAM (Sequence Alignment/Map) is a text-based format for storing read alignments — the results of mapping short sequencing reads to a reference genome using tools like BWA-MEM, HISAT2, STAR, or Bowtie2. BAM is the binary compressed equivalent of SAM, dramatically smaller in size and much faster to parse. BAM files are the primary storage format for aligned sequencing data in production pipelines.
The CIGAR (Compact Idiosyncratic Gapped Alignment Report) string describes the alignment of a read to the reference:
| Op | Meaning | Example |
|---|---|---|
M | Match or mismatch to reference (consumes both) | 151M = 151 bp aligned |
I | Insertion relative to reference (read only) | 10M2I139M = 10 match, 2 bp insertion, 139 match |
D | Deletion from reference (reference only) | 50M3D98M = 50 match, 3 bp deletion, 98 match |
S | Soft clip — read bases not aligned (at ends) | 10S141M = 10 bp clipped, 141 aligned |
N | Skipped region (intron in RNA-seq alignments) | 50M1000N50M = exon-intron-exon junction |
| Format | Type | Compression | Typical Size (30× WGS) | Use Case |
|---|---|---|---|---|
| SAM | Text | None | ~300 GB | Human inspection, small files, pipe input/output |
| BAM | Binary | BGZF (gzip-based) | ~100 GB | Standard storage and processing |
| CRAM | Binary | Reference-based | ~40–60 GB | Long-term archiving, reference-based compression |
| Conversion | Tool | Command Example |
|---|---|---|
| SAM → BAM | samtools | samtools view -bS input.sam > output.bam |
| BAM → CRAM | samtools | samtools view -C -T ref.fa input.bam > output.cram |
| FASTQ → FASTA | seqtk | seqtk seq -a input.fastq.gz > output.fasta |
| VCF → BED | bedtools | bedtools intersect -a variants.vcf -b regions.bed |
| GFF → BED | BEDOPS | gff2bed < annotation.gff > annotation.bed |
| BAM sort + index | samtools | samtools sort input.bam -o sorted.bam && samtools index sorted.bam |
| Format | Purpose | Text/Binary | Indexed? | Typical Extension |
|---|---|---|---|---|
| FASTA | Reference sequences | Text | .fai (samtools) | .fa / .fasta |
| FASTQ | Raw sequencing reads | Text (often gzipped) | No | .fastq / .fq / .fastq.gz |
| VCF | Genetic variants | Text (bgzipped) | .tbi / .csi (tabix) | .vcf / .vcf.gz / .bcf |
| BED | Genomic intervals | Text | .bai (UCSC) / .tbi (tabix) | .bed / .bed.gz |
| SAM | Read alignments (text) | Text | No | .sam |
| BAM | Read alignments (binary) | Binary (BGZF) | .bai (samtools) | .bam |
| CRAM | Compressed alignments | Binary | .crai | .cram |
| GFF/GTF | Gene annotations | Text | .tbi (tabix) | .gff / .gtf / .gff3 |
Subscribe for updates on new tools, format specifications, and analysis tutorials.