✕
CN EN

Dimensionality reduction & clustering / UMAP

Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.

How to read it

UMAP

UMAP places samples that are close in high-dimensional space together and is widely used to show clusters in single-cell data; compared with t-SNE it keeps somewhat more large-scale structure and runs faster.

Larger n_neighbors emphasises global structure, smaller values local detail (5–50 is typical); min_dist controls how tightly points pack within clusters (0.0–0.5 is typical).

Cluster sizes, distances between clusters and coordinate values are still not directly comparable. With the same random seed the result is reproducible; different seeds give different layouts.

For gene-expression matrices log-transform first and reduce to the top 30–50 principal components (default 50). Trustworthiness (0–1, higher is better) measures how well local neighbours are preserved.

This is a simplified in-browser implementation (exact neighbours, spectral initialisation, single-threaded optimisation): the structure is similar to the widely used Python implementation but coordinates are not identical point by point. Set the parameters, then press “Run”.

Method

McInnes L., Healy J. & Melville J. (2018) arXiv:1802.03426. The model follows umap-learn (a and b fitted from min_dist and spread, within 1e-4 of umap-learn find_ab_params); differences: exact k-nearest neighbours instead of NN-descent, own random sequence, single thread. On the example data trustworthiness is within 0.03 of umap-learn.

Data size

Up to 5,000 samples.

Need a full analysis?

Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.