Sequence utilities / Motif search
Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.
Finds every site matching the motif. Nucleotide motifs use IUPAC codes (R = A/G, Y = C/T, N = any, etc.). Both strands are searched by default, and minus-strand hits are reported in plus-strand coordinates (1-based, inclusive).
The chart shows where hits fall along the sequence (or hits per sequence when there are several), so clusters stand out.
“Observed/expected” compares the hit count with the number expected from the base composition of your sequences: well above 1 means the motif is enriched (e.g. Chi sites in the E. coli genome); close to 1 is what chance would give. Palindromic motifs (such as GATC) count once even though both strands match.
Choose “Regular expression” for more flexible patterns; these are matched as written, without IUPAC expansion or an expected value.
IUPAC-IUB ambiguity codes (Cornish-Bowden 1985); expected counts use an independent-base (zero-order Markov) model. On Chi sites see Smith 2012 (Microbiol Mol Biol Rev 76:217).
Up to 50 Mb per sequence; at most 200,000 hits are kept.
Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.