Dimensionality reduction & clustering / PCA
Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.
Each point is a sample; samples close together differ little across all variables. The percentage in each axis title is the share of total variance explained by that component.
If variables have different units (e.g. concentrations and pH) or very different magnitudes, tick “Scale to unit variance” (correlation-matrix PCA). Log-transform counts or expression values first.
Ellipses are 95% normal confidence ellipses of each group’s scores (as ggplot2 stat_ellipse), drawn for groups with at least 4 samples. They describe within-group spread and are not a test; use PERMANOVA to test group differences.
In the biplot, arrow direction shows how a variable relates to the two components and longer arrows contribute more; arrows are rescaled to the score range, so read direction and relative length only.
Component signs are arbitrary (here the largest-magnitude loading of each component is made positive), so a plot mirrored relative to other software is expected.
Jolliffe I. T. (2002) Principal Component Analysis, 2nd ed. (Springer). Results match scikit-learn PCA (scaling uses the n − 1 standard deviation); ellipses are normal-theory confidence ellipses with an F-based radius (Fox & Weisberg 2011).
Up to about 50,000 samples or variables (large matrices take a few seconds).
Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.