Machine learning / Regression (linear / ridge / random forest)
Your data stays in your browser: files are read and processed on your device, never uploaded to any server, and cleared when you close the page. Free, no sign-in.
R², RMSE and MAE come from k-fold cross-validation: each fold is predicted by a model trained on the other samples, and the fold values are averaged. R² closer to 1 is better and can be negative (worse than predicting the mean).
The scatter plot shows observed values against cross-validated predictions; points close to the diagonal mean accurate predictions, while a horizontal band means the model explains little.
For linear and ridge regression with standardised features, a coefficient is the change in outcome per standard deviation of the feature, so sizes are comparable; coefficients in original units are also given. A larger ridge α shrinks coefficients more, which helps with many or correlated features.
Random forests capture non-linearity and interactions; impurity importance favours features with many distinct values, so compare it with permutation importance (the fall in R² when a feature is shuffled).
By default folds are cut in row order without shuffling (as in common implementations); if the table is sorted by outcome, tick “Shuffle before splitting”.
Linear and ridge regression are solved exactly via the normal equations, with the same objective as scikit-learn LinearRegression / Ridge (unpenalised intercept); k-fold splits match scikit-learn KFold; random forest follows Breiman (2001) Machine Learning 45:5–32; R² is the coefficient of determination 1 − SSE/SST.
Runs in your browser: up to about 5,000 samples and 300 features is recommended; random forests are slower, so reduce the number of trees if needed.
Send us your data and research question and you will receive a written plan within 1 working day: analysis steps, parameter rationale, deliverables and timeline. Quoted per project.