✕
CN EN

Machine learning / Clustering evaluation

Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.

How to read it

Clustering evaluation

The silhouette compares how close each sample is to its own cluster with the nearest other cluster: near 1 means a clear assignment, near 0 means between two clusters, and negative values suggest a likely misassignment. The silhouette plot is grouped by cluster; clusters with many negative values deserve a second look.

The Calinski–Harabasz index (between- to within-cluster dispersion) is better when higher, and the Davies–Bouldin index (average similarity of each cluster to its most similar one) is better when lower; both are for comparing clusterings of the same data.

If known classes are supplied, ARI (adjusted Rand index, about 0 for random grouping) and NMI (normalised mutual information, 0–1) are reported with a cluster × class contingency table, showing which classes were merged or split.

Distances are Euclidean; keep “Standardise” on when features have different units. Without a cluster column, k-means is run for k = 2 up to the chosen maximum and a k is suggested by mean silhouette — a guide rather than a verdict.

The PCA scatter shows the clusters on the first two principal components; it is a projection, so overlap there does not necessarily mean overlap in the full space.

Method

Silhouette: Rousseeuw (1987) J Comput Appl Math 20:53–65; CH: Caliński & Harabasz (1974) Commun Stat 3:1–27; DB: Davies & Bouldin (1979) IEEE TPAMI 1:224–227; ARI: Hubert & Arabie (1985) J Classif 2:193–218; NMI with arithmetic normalisation. Definitions and edge cases match scikit-learn.

Data size

The silhouette needs all pairwise distances, so up to 5,000 samples.

Need a full analysis?

Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.