✕
CN EN

Variants & FASTQ / FASTQ / FASTA conversion and subsampling

Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.

How to read it

FASTQ / FASTA conversion and subsampling

Order of steps: fixed trimming at both ends, then 3′ quality trimming (bases below the threshold are removed one by one until a base reaches it), then filtering by length, mean quality and N count, and finally random subsampling and output.

Subsampling by fraction keeps each read independently with that probability; by number of reads it draws exactly N reads by reservoir sampling and keeps the original order. A fixed seed gives the same result each time; for paired-end data use the same settings and seed on both files and check the pairing after filtering.

The quality encoding (Phred+33 or Phred+64) is detected automatically; FASTA headers keep the original ID and description.

Copy or download the output text; the histogram compares read lengths before and after, and the table lists how many reads each step removed.

Method

Trimming and filtering are standard read pre-processing steps (cf. Chen et al. 2018, Bioinformatics 34:i884); subsampling uses reservoir sampling (Vitter 1985, ACM TOMS 11:37) with a seeded generator, and the output matches an independent Python reference byte for byte.

Data size

Runs in your browser; uncompressed FASTQ up to about 200 MB. Compressed .gz files cannot be read directly yet.

Need a full analysis?

Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.