Genomic & Bioinformatics File Formats

A technical reference for the major file formats used in genomics, sequence analysis, and bioinformatics pipelines — with format specifications, examples, and usage guidance.

Bioinformatics relies on standardized file formats that allow genomic data to be processed consistently across different software tools, pipelines, and institutions. Understanding these formats is essential for working with sequencing data, variant calls, genome annotations, and alignment results. This reference covers the six most important formats you will encounter in modern genomics workflows, with annotated examples for each.

FASTA Format

Sequence storage — DNA, RNA, and protein sequences

FASTA is the simplest and most universally used format for storing biological sequences. Named for the FASTA sequence alignment program (1985), it stores one or more sequences with associated metadata in a plain text format that is both human-readable and machine-parsable.

Format Structure

Each sequence entry consists of:

  1. Header line: Begins with > (greater-than symbol), followed by a sequence identifier and optional description
  2. Sequence data: One or more lines of sequence characters. Lines are typically 60–80 characters wide by convention.
Example — Multiple Sequences
# FASTA format example with three sequences >ENSG00000012048.23 BRCA1 gene, chromosome 17, forward strand ATGGATTTATCTGCTCTTCGCGTTGAAGAAGTACAAAATGTCATTAATGCTATGCAGAAA ATCTTAGAGTGTCCCATCTGTCTGGAGTTGATCAAGGAACCTGTCTCCACAAAGTGTGAC CACAGCTTTCTGGAGGATGCAAACACTAAGCAAGGTAGTGCCATTGTGGTCAGCCTTTAC >NM_007294.4 BRCA1 mRNA transcript, RefSeq ATGGATTTATCTGCTCTTCGCGTTGAAGAAGTACAAAATGTCATTAATGCTATGCAGAAA >P38398.3 BRCA1 protein sequence, UniProt MDLSALRVEEVQNVINAMQKILECPICLELIKEPVSTKCDHIFCKFCMLKLLNQKKGPSQ CPLCKNDITKRSLQESTRFSQLVEELLKIICAFQLDTGLEYANSYNFAKKENNSPEHLKD
Nucleotide IUPAC Codes
CodeBases RepresentedMeaning
AAAdenine
CCCytosine
GGGuanine
TTThymine (DNA) / Uracil (RNA)
NA, C, G, TAny nucleotide (unknown)
RA, GPurine
YC, TPyrimidine
WA, TWeak
SG, CStrong
Common Uses
  • Reference genome sequences (human GRCh38, mouse GRCm39)
  • Protein sequence databases (UniProt, NCBI RefSeq proteins)
  • Transcript sequences (GENCODE, Ensembl)
  • BLAST and sequence alignment queries
  • Primer sequences, probe sequences

FASTQ Format

Raw sequencing reads with per-base quality scores

FASTQ extends FASTA by including per-base quality scores alongside the sequence. It is the standard output format of next-generation sequencing (NGS) instruments — Illumina, Oxford Nanopore, PacBio platforms all produce FASTQ output. Every short read sequencing experiment begins with FASTQ files.

Format Structure

Each read is represented by exactly 4 lines:

  1. Line 1 — Identifier: Begins with @, followed by sequence identifier and optional description (instrument, flowcell, coordinates)
  2. Line 2 — Sequence: Raw nucleotide sequence
  3. Line 3 — Separator: Always + (optionally followed by the identifier again)
  4. Line 4 — Quality scores: ASCII-encoded Phred quality scores, one character per base
Example
# FASTQ format — 4 lines per read @NB501234:1:H3GVFBGX7:1:11101:8134:1049 1:N:0:ATCACGAT GCTTACGAATTTCGATCGATCGTAGCTAGCTAGCTAGCTTTTACGTACGATCGAT + AAAAAEEEE6EEAEEEEEEEEEEEEEEEEEEEAEEEEEEEEEEEEEEEEEEEEEEE @NB501234:1:H3GVFBGX7:1:11101:15917:1049 1:N:0:ATCACGAT ATCGATCGATCGATCGATCGTTTTACGTACGTACGATCGATCGATCGATCGATCG + 6AAEEEEEE/EEEEEEEEAEEE6/AEEEEEEEAEE//EEEEAEEEE<EEAEAEEA
Phred Quality Scores

Quality scores (Q-scores) quantify the probability of a base call error. The relationship is: Q = -10 × log₁₀(P), where P is the error probability.

Q-ScoreError ProbabilityAccuracyASCII Character (Phred+33)
101 in 1090%+ (ASCII 43)
201 in 10099%5 (ASCII 53)
301 in 1,00099.9%? (ASCII 63)
401 in 10,00099.99%I (ASCII 73)
40+ (Q41)<1 in 10,000>99.99%J (ASCII 74)

Modern Illumina sequencing typically achieves Q30+ for 75–90% of bases. Reads are typically quality-filtered (remove low-quality bases/reads) and adapter-trimmed before downstream analysis using tools like Trimmomatic, fastp, or Cutadapt.

Paired-End Sequencing

Most Illumina sequencing is paired-end: each fragment is sequenced from both ends, producing two FASTQ files (R1 and R2) where reads at the same position in each file are paired. Paired reads are more informative for alignment and structural variant detection.

VCF — Variant Call Format

Genetic variants: SNPs, indels, structural variants

VCF (Variant Call Format) is the standard format for storing genetic variants identified from sequencing data — single nucleotide polymorphisms (SNPs), insertions/deletions (indels), and structural variants. It is produced by variant calling tools such as GATK HaplotypeCaller, DeepVariant, and Strelka2, and used by major variant databases (dbSNP, ClinVar, gnomAD).

Format Structure
##fileformat=VCFv4.3 ##fileDate=20240115 ##reference=GRCh38/hg38 ##FILTER=<ID=PASS,Description="All filters passed"> ##INFO=<ID=DP,Number=1,Type=Integer,Description="Total Depth"> ##INFO=<ID=AF,Number=A,Type=Float,Description="Allele Frequency"> ##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype"> ##FORMAT=<ID=GQ,Number=1,Type=Integer,Description="Genotype Quality"> ##FORMAT=<ID=AD,Number=R,Type=Integer,Description="Allelic Depth"> ##FORMAT=<ID=DP,Number=1,Type=Integer,Description="Read Depth"> #CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1 chr17 41197694 rs80357914 G A 50 PASS DP=45;AF=0.489 GT:GQ:AD:DP 0/1:99:22,23:45 chr17 41244000 rs28897686 C T . PASS DP=32;AF=0.5 GT:GQ:AD:DP 0/1:85:16,16:32 chr17 41246506 . AG A 200 PASS DP=60;AF=0.35 GT:GQ:AD:DP 0/1:99:39,21:60
Core VCF Columns
ColumnDescription
CHROMChromosome (chr1, chrX, chr17...)
POS1-based position of the first base of the variant
IDdbSNP rs identifier, or "." if unknown
REFReference allele at this position
ALTAlternate allele(s) — comma-separated for multi-allelic sites
QUALPhred-scaled quality score for ALT allele call
FILTERPASS if variant passes all filters; filter name if failed
INFOSemicolon-separated key=value annotations (depth, AF, functional impact)
FORMATColon-separated list of per-sample format fields (GT, GQ, DP, AD)
SAMPLEOne column per sample with colon-separated values matching FORMAT
Genotype Encoding (GT field)
  • 0/0 — Homozygous reference (no variant)
  • 0/1 — Heterozygous (one reference, one alternate allele)
  • 1/1 — Homozygous alternate
  • 0|1 — Phased heterozygous (paternal|maternal allele)
  • ./. — Missing data (insufficient coverage)

BED Format

Genomic intervals and regions — gene locations, peaks, regulatory elements

BED (Browser Extensible Data) format is the standard way to represent genomic intervals — specific chromosomal regions defined by start and end coordinates. It is used for gene annotations, ChIP-seq peaks, ATAC-seq open chromatin regions, CpG islands, repeat elements, regulatory regions, and virtually any genomic feature defined by position.

# BED format — tab-delimited, 0-based half-open coordinates # Fields: chrom, chromStart, chromEnd [name, score, strand, thickStart, thickEnd, itemRgb, blockCount, blockSizes, blockStarts] chr17 41196311 41277500 BRCA1 0 + chr13 32889616 32974405 BRCA2 0 + chr7 55019016 55211628 EGFR 0 + chrX 153701385 153702385 ChIP_peak_1 87 .
Important — 0-based coordinates: BED uses a 0-based, half-open coordinate system. Position 0 is the first base of the chromosome. The interval [start, end) includes the start base but excludes the end base. This differs from VCF (1-based) and GFF (1-based, inclusive) — be careful when converting between formats or doing arithmetic with coordinates.
BED Core Fields
FieldRequiredDescription
chromYesChromosome name (chr1, chrX)
chromStartYesStart position, 0-based
chromEndYesEnd position, exclusive (0-based)
nameNoFeature name/identifier
scoreNoScore 0–1000 (peak signal, conservation score)
strandNo+ (forward), - (reverse), or . (unknown)

BAM/SAM Format

Sequence alignment/map — aligned reads against a reference genome

SAM (Sequence Alignment/Map) is a text-based format for storing read alignments — the results of mapping short sequencing reads to a reference genome using tools like BWA-MEM, HISAT2, STAR, or Bowtie2. BAM is the binary compressed equivalent of SAM, dramatically smaller in size and much faster to parse. BAM files are the primary storage format for aligned sequencing data in production pipelines.

SAM File Structure
# SAM header lines (begin with @) @HD VN:1.6 SO:coordinate @SQ SN:chr17 LN:83257441 @RG ID:SampleA SM:Patient001 PL:ILLUMINA LB:WGS_lib1 @PG ID:bwa PN:bwa VN:0.7.17 CL:bwa mem -M GRCh38.fa R1.fastq R2.fastq # Alignment records (one per read) # QNAME FLAG RNAME POS MAPQ CIGAR RNEXT PNEXT TLEN SEQ QUAL read_1 99 chr17 41197694 60 151M = 41197845 302 ATCGATCG... AAAEEEEE... read_2 147 chr17 41197845 60 10S141M = 41197694 -302 CGTAGCAT... EEEAAAEE...
CIGAR String — Alignment Encoding

The CIGAR (Compact Idiosyncratic Gapped Alignment Report) string describes the alignment of a read to the reference:

OpMeaningExample
MMatch or mismatch to reference (consumes both)151M = 151 bp aligned
IInsertion relative to reference (read only)10M2I139M = 10 match, 2 bp insertion, 139 match
DDeletion from reference (reference only)50M3D98M = 50 match, 3 bp deletion, 98 match
SSoft clip — read bases not aligned (at ends)10S141M = 10 bp clipped, 141 aligned
NSkipped region (intron in RNA-seq alignments)50M1000N50M = exon-intron-exon junction
BAM vs SAM vs CRAM
FormatTypeCompressionTypical Size (30× WGS)Use Case
SAMTextNone~300 GBHuman inspection, small files, pipe input/output
BAMBinaryBGZF (gzip-based)~100 GBStandard storage and processing
CRAMBinaryReference-based~40–60 GBLong-term archiving, reference-based compression

Format Conversion Tools

ConversionToolCommand Example
SAM → BAMsamtoolssamtools view -bS input.sam > output.bam
BAM → CRAMsamtoolssamtools view -C -T ref.fa input.bam > output.cram
FASTQ → FASTAseqtkseqtk seq -a input.fastq.gz > output.fasta
VCF → BEDbedtoolsbedtools intersect -a variants.vcf -b regions.bed
GFF → BEDBEDOPSgff2bed < annotation.gff > annotation.bed
BAM sort + indexsamtoolssamtools sort input.bam -o sorted.bam && samtools index sorted.bam

Format Comparison Summary

FormatPurposeText/BinaryIndexed?Typical Extension
FASTAReference sequencesText.fai (samtools).fa / .fasta
FASTQRaw sequencing readsText (often gzipped)No.fastq / .fq / .fastq.gz
VCFGenetic variantsText (bgzipped).tbi / .csi (tabix).vcf / .vcf.gz / .bcf
BEDGenomic intervalsText.bai (UCSC) / .tbi (tabix).bed / .bed.gz
SAMRead alignments (text)TextNo.sam
BAMRead alignments (binary)Binary (BGZF).bai (samtools).bam
CRAMCompressed alignmentsBinary.crai.cram
GFF/GTFGene annotationsText.tbi (tabix).gff / .gtf / .gff3

Stay Updated on Bioinformatics Resources

Subscribe for updates on new tools, format specifications, and analysis tutorials.