✕
CN EN

Table utilities / Missing-value handling

Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.

How to read it

Missing-value handling

Columns and rows with too many missing values (above 50% by default) are dropped first, then remaining gaps are filled. Blank, NA, NaN, null and “-” all count as missing.

Mean or median filling is simple but shrinks the variance; the minimum or half-minimum suits values missing below the detection limit in metabolomics/proteomics; kNN fills with the average of the k most similar rows (Euclidean distance on the other columns) and is usually closer to the truth.

kNN matches the KNNImputer of common machine-learning libraries: distances use only columns observed in both rows, scaled up proportionally, and neighbours are chosen among rows that have the column observed.

“Imputed cells” lists every filled position and value for checking; if imputed data are used for statistical tests, state the method in your report.

Method

kNN imputation: Troyanskaya et al. 2001 (Bioinformatics 17:520); distance as in scikit-learn KNNImputer (nan_euclidean).

Data size

Up to 100,000 rows; kNN imputation is recommended for up to about 5,000 rows.

Need a full analysis?

Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.