Variants & FASTQ / FASTQ / FASTA conversion and subsampling
Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.
Order of steps: fixed trimming at both ends, then 3′ quality trimming (bases below the threshold are removed one by one until a base reaches it), then filtering by length, mean quality and N count, and finally random subsampling and output.
Subsampling by fraction keeps each read independently with that probability; by number of reads it draws exactly N reads by reservoir sampling and keeps the original order. A fixed seed gives the same result each time; for paired-end data use the same settings and seed on both files and check the pairing after filtering.
The quality encoding (Phred+33 or Phred+64) is detected automatically; FASTA headers keep the original ID and description.
Copy or download the output text; the histogram compares read lengths before and after, and the table lists how many reads each step removed.
Trimming and filtering are standard read pre-processing steps (cf. Chen et al. 2018, Bioinformatics 34:i884); subsampling uses reservoir sampling (Vitter 1985, ACM TOMS 11:37) with a seeded generator, and the output matches an independent Python reference byte for byte.
Runs in your browser; uncompressed FASTQ up to about 200 MB. Compressed .gz files cannot be read directly yet.
Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.