✕
CN EN

Sequence utilities / Sequence logo

Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.

How to read it

Sequence logo

Each column is a position; the height of the letter stack is its information content in bits (up to 2 bits for DNA and about 4.32 bits for proteins) — taller means more conserved. Each letter’s height is proportional to its frequency, with the most common on top.

With “Probability” on the y axis every column has height 1, showing composition only rather than conservation.

Information = log2(alphabet size) − Shannon entropy. With few sequences entropy is underestimated; the table can report the small-sample-corrected value, while the chart uses observed frequencies without correction.

The input must be equal-length, aligned sites (e.g. anchored on the start codon or transcription start); in the consensus, upper case marks positions with ≥1 bit.

Method

Sequence logos: Schneider & Stephens 1990 (Nucleic Acids Res 18:6097); small-sample correction as in Crooks et al. 2004 (WebLogo, Genome Res 14:1188).

Data size

Up to 300 positions.

Need a full analysis?

Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.