library(rtemis.draw)10 Variable importance
Use draw_varimp() to rank predictors and compare their importance across resamples. Use rtemis::plot_varimp() for importance stored in a fitted rtemis object. These views build on bars and boxplots.
10.1 Rank predictors
A named numeric vector supplies one score per variable. These illustrative permutation scores measure the increase in prediction error after shuffling each predictor:
scores <- c(Weight = 0.84, Power = 0.62, Displacement = 0.48,
Gearing = 0.31, Cylinders = 0.22, Carburetors = 0.14)
importance_chart <- draw_varimp(scores, top_n = NULL, xlab = "Increase in RMSE",
title = "Permutation importance")
importance_chartBy default, variables are ranked by score magnitude, with the largest at the top. Use top_n to focus on the leading predictors. Scores keep the meaning and units of their source: permutation importance, split gain, and regression coefficients describe different quantities.
10.1.1 Use thin segments
A small bar_width produces a lighter zero-to-score display. Thickness is in pixels; a horizontal layout keeps predictor names easy to read:
draw_varimp(scores, bar_width = 4, xlab = "Increase in RMSE",
title = "Predictor importance")10.2 Choose a ranking direction
Tables use a variable column and one or more numeric measures. Select a measure by name. When smaller values are preferred, rank by signed value and set decreasing = FALSE:
losses <- data.frame(
variable = names(scores),
validation_rmse = c(2.1, 2.4, 2.7, 3.0, 3.2, 3.5)
)
draw_varimp(losses, measure = "validation_rmse", rank_by = "signed",
decreasing = FALSE, top_n = 4, xlab = "Validation RMSE",
title = "Best single-predictor models")These illustrative losses come from separate single-predictor models. For signed coefficients, rank_by = "magnitude" highlights the largest absolute effects while retaining their signs in the plot.
10.3 Summarize repeated fits
Add a fold column when you have one importance score per predictor and resample. The following illustrative scores vary around the predictor ranking above:
set.seed(36)
fold_scores <- expand.grid(variable = names(scores), fold = paste0("Fold", 1:12),
stringsAsFactors = FALSE)
fold_scores[["rmse_increase"]] <- rep(unname(scores), 12) *
rlnorm(nrow(fold_scores), sdlog = 0.2)The default summary is the mean across available folds. Use summary = "median" for a median ranking:
draw_varimp(fold_scores, measure = "rmse_increase", xlab = "Increase in RMSE",
title = "Mean permutation importance")10.3.1 Show the distribution
type = "boxplot" shows the variation across folds. The summary still controls the ordering, while every available fold score appears as a point:
draw_varimp(fold_scores, measure = "rmse_increase", type = "boxplot", xlab = "Increase in RMSE",
title = "Importance across resamples")Hover points to see their fold IDs. Use whisker = 0 for full-range whiskers, or boxpoints = "outliers" to reduce the number of displayed points. These plots are useful for distinguishing a consistently strong predictor from one whose importance varies between fits.
10.4 Fitted rtemis objects
For a fitted or resampled rtemis result, plot_varimp() extracts the stored scores. This example fits classification trees across five iris folds and shows the distribution of their split importance. It additionally requires rtemis:
resampled <- rtemis::train(
iris,
hyperparameters = rtemis::setup_CART(),
outer_resampling_config = rtemis::setup_KFold(n_resamples = 5L, seed = 31L),
execution_config = rtemis::setup_SerialExecution(), verbosity = 0L
)
rtemis::plot_varimp(resampled, type = "boxplot", whisker = 0)Qualify the plotting generic when both packages are loaded. For an ordinary fitted model, use rtemis::plot_varimp(model) to show its importance ranking. See Plotting model objects for prediction, metric, and execution views of fitted objects.
10.5 Save a figure
Export the chart as an SVG for a report or presentation:
save_drawing(importance_chart, "importance.svg", width = 800, height = 500)See Export for output options and Chart configs for reusable plotting settings.