Sequence utilities / Sequence length statistics
Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.
The summary gives the number of sequences, total length, N50, mean length and pooled GC%. The overall table adds min/max, median, L50, N75, N90 and the proportion of N bases.
N50 is the length of the sequence at which, adding sequences from longest to shortest, the running total first reaches half of the total length; L50 is how many sequences that took. A larger N50 means a more contiguous assembly.
GC% counts only A/C/G/T(U)/S/W; N and other ambiguity codes are left out of the denominator, as in common tools.
The histograms show the distribution of lengths (and per-sequence GC%); adjust the number of bins as needed. Download the per-sequence table for filtering or plotting. For protein sequences only lengths are reported.
N50/L50 follow the standard assembly-assessment definitions (Gurevich et al. 2013, Bioinformatics 29:1072, QUAST); GC semantics match Biopython gc_fraction.
Runs in your browser; up to roughly 100 MB of sequence is recommended.
Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.