Genomic intervals & annotation / Annotation statistics and validation
Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.
The summary gives the numbers of genes, transcripts and unique exons; the distribution table lists minimum, median, mean and maximum of gene span, spliced transcript length, exon length, exons per transcript and transcripts per gene.
Validation reports non-integer coordinates or end before start, invalid strands, missing gene_id or gene rows in GTF, undefined Parent in GFF3, exons outside their transcript or gene, duplicated or overlapping exons within a transcript, transcripts without exons, and children on a different chromosome or strand from their gene; the first 10 line numbers of each are shown.
GTF is grouped by gene_id / transcript_id; GFF3 by the ID / Parent hierarchy (CDS attached directly to genes, as in bacterial annotations, are recognised).
Charts show feature-type counts, the chosen length distribution and gene biotypes; the per-gene and per-transcript tables can be downloaded for further filtering.
Parsing and grouping follow GTF2.2 and GFF3 (Sequence Ontology specification 1.26); counts and length distributions were checked against an independent Python reference implementation.
Runs in your browser; annotation files up to roughly 100 MB (about 300,000 lines) are recommended. For a complete human GTF (about 1.5 GB uncompressed), extract the chromosomes or genes you need first.
Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.