library(rtemis.draw)7 Classification diagnostics
Use draw_confusion() to see which classes a model confuses, draw_roc() to compare its ranking performance across decision thresholds, and draw_calibration() to inspect its predicted probabilities. The examples below use held-out iris observations and ordinary R model predictions.
7.1 Prepare predictions
We will predict whether a flower is Iris virginica from its sepal measurements. Reserve 15 observations from each species for testing, then fit a logistic regression on the remaining observations:
set.seed(42)
training_rows <- unlist(lapply(
split(seq_len(nrow(iris)), iris[["Species"]]), sample, size = 35
))
training <- iris[training_rows, ]
test <- iris[-training_rows, ]
training[["virginica"]] <- training[["Species"]] == "virginica"
sepal_model <- glm(virginica ~ Sepal.Length + Sepal.Width,
data = training, family = binomial())
probability <- predict(sepal_model, newdata = test, type = "response")
reference <- ifelse(test[["Species"]] == "virginica", "Virginica", "Other")
predicted <- ifelse(probability >= 0.5, "Virginica", "Other")7.2 Confusion matrices
Pass reference and predicted labels to draw_confusion(). Here predictions use a probability threshold of 0.5:
confusion_chart <- draw_confusion(
reference, predicted, classes = c("Virginica", "Other"), font_size = 11
)
confusion_chartRows show reference classes and columns show predicted classes. Labels are counts; color intensity shows the fraction within each reference row. Teal highlights correct predictions and magenta highlights errors. Hover a cell for its count and row fraction.
The surrounding summaries report sensitivity (recall), specificity, positive predictive value (precision), and negative predictive value for each class. Accuracy is the fraction of all predictions that are correct; balanced accuracy (BA) is the mean recall across classes.
Rates in cells and hover text use two decimal places. Set digits to change this precision; counts remain integers.
7.3 ROC curves
Pass the same reference labels and predicted probabilities to draw_roc(). positive identifies the class represented by the probability vector:
roc_chart <- draw_roc(reference, probability, positive = "Virginica")
roc_chartEach point on the curve represents a decision threshold. The horizontal axis is the false-positive rate and the vertical axis is sensitivity. The legend reports area under the curve (AUC); the dashed diagonal shows chance ranking. Higher scores must indicate stronger evidence for the positive class.
The legend sits above the plotting area. Set legend_position = "top-right" to right-align it in the same band. Add legend_placement = "inside" to place it over the data, or use legend = FALSE to hide it. For example:
draw_roc(reference, probability, positive = "Virginica",
legend_position = "top-right")7.3.1 Compare models
Named lists let you compare models on the same observations. Fit a second model using petal width, then draw both ROC curves together:
petal_model <- glm(virginica ~ Petal.Width,
data = training, family = binomial())
petal_probability <- predict(petal_model, newdata = test, type = "response")
draw_roc(
list(Sepal = reference, Petal = reference),
list(Sepal = probability, Petal = petal_probability),
positive = "Virginica"
)Click a legend entry to hide or restore a model’s curve. The same list syntax works for predictions from separate training and test samples.
7.4 Probability calibration
A model can rank observations well while assigning probabilities that are too high or too low. Compare mean predicted probability with the observed fraction of positive cases using the same held-out predictions:
calibration_chart <- draw_calibration(
list(Sepal = reference, Petal = reference),
list(Sepal = probability, Petal = petal_probability),
positive = "Virginica", n_bins = 5L
)
calibration_chartEach point summarizes a bin of predicted probabilities. The dashed diagonal represents perfect agreement. Quantile bins are computed separately for each model; tied probabilities stay together, so fewer than five bins can remain. Set bin_method = "equidistant" to use equal-width intervals from zero to one. With only 45 held-out observations, these curves are descriptive and can be noisy; they are not confidence intervals or a test of calibration.
The legend reports the Brier score, the mean squared error of the individual probabilities. Lower values indicate better probability predictions; the score reflects more than calibration alone. Hover a point for its bin count, mean probability, observed proportion, and boundaries. Small ticks along the lower edge show individual probabilities. A legend click toggles a model’s curve and its probability rug together; rug = FALSE hides the rugs.
7.5 Multiclass predictions
For several classes, supply a probability matrix with one named column per class. This example uses linear discriminant analysis from MASS, with sepal length and width as predictors:
multiclass_model <- MASS::lda(Species ~ Sepal.Length + Sepal.Width,
data = training)
multiclass_prediction <- predict(multiclass_model, newdata = test)
draw_roc(test[["Species"]], multiclass_prediction[["posterior"]])Each curve compares one species with the other two. Its AUC describes that one-versus-rest ranking.
A table of reference and predicted classes gives a compact overview of the same model. show_metrics = FALSE emphasizes the class-to-class pattern:
counts <- table(Reference = test[["Species"]],
Predicted = multiclass_prediction[["class"]])
draw_confusion(counts, show_metrics = FALSE, font_size = 11)7.6 Save a figure
Save either chart as an SVG for a report or presentation:
save_drawing(confusion_chart, "confusion.svg", width = 800, height = 550)
save_drawing(roc_chart, "roc.svg", width = 650, height = 650)
save_drawing(calibration_chart, "calibration.svg", width = 650, height = 650)See Export for output options and Chart configs for reusable settings. For fitted rtemis objects, use rtemis::plot_true_pred() and rtemis::plot_roc(); see Plotting model objects.