✕
CN EN

Sequence utilities / Sequence formatting

Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.

How to read it

Sequence formatting

Brings FASTA files from different sources to one layout: fixed line width, consistent case, tidy headers, and optional de-duplication by sequence or ID.

“ID only” drops the description after the first space; “special characters replaced by _” suits software with strict ID rules (e.g. tree-building programs); “Rename” keeps the original ID in the description for tracing.

De-duplication by sequence ignores case and keeps the first occurrence; removed records and the record they duplicate are listed in the table below.

Order of operations: remove non-letters → remove gaps → change case → drop empty → de-duplicate.

Method

FASTA format (Pearson & Lipman 1988, PNAS 85:2444).

Data size

Runs in your browser; up to roughly 100 MB of sequence is recommended.

Need a full analysis?

Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.