Bind binary observation records to a reliability diagram. Each row contains an observed outcome (zero or one) and its predicted probability, optionally grouped by sample or model. Configuration carries column names, not data.
Usage
CalibrationConfig(
dat_path = NULL,
title = NULL,
origin = NULL,
writer = NULL,
legend_position = "top",
legend_placement = "outside",
observed = "observed",
probability = "probability",
group = NULL,
n_bins = 10L,
bin_method = "quantile",
na_rm = TRUE,
mode = "lines+markers",
show_brier = TRUE,
rug = TRUE,
rug_size = 8,
rug_opacity = 0.3,
point_size = 7,
line_width = 2,
digits = 3L,
diagonal = TRUE,
diagonal_color = "#888888",
palette = NULL,
legend = TRUE,
square = TRUE,
xlab = "Mean predicted probability",
ylab = "Observed proportion"
)Arguments
- dat_path
Optional Character: Path to the data, read at draw time. The serializable alternative to passing
datatodraw().- title
Optional Character: Chart title.
- origin
Optional Named character {"user", "default", "derived"}: Where each value came from, one entry per settable property. Absent on an authored config; written by the interface that resolved it.
- writer
Optional Named character: Which interface wrote the config, as
nameandversion. Absent on an authored config.- legend_position
Character {"top", "bottom", "left", "right", "top-left", "top-right", "bottom-left", "bottom-right"}: Legend anchor. Top/bottom anchors use horizontal rows; left/right anchors use a vertical column. Corner anchors align within the top or bottom row.
- legend_placement
Character {"outside", "inside"}: Relation to the plotting area. Outside placement reserves space for the complete legend; inside placement overlays the data. Neither setting adds a missing legend.
- observed
Character: Column containing binary outcomes, zero or one.
- probability
Character: Column containing predicted probabilities.
- group
Optional Character: Column identifying samples or models.
- n_bins
Integer
[1, Inf): Requested number of bins per group.- bin_method
Character {"quantile", "equidistant"}: Binning rule.
- na_rm
Logical: Remove incomplete observation/probability pairs.
- mode
Character {"lines", "markers", "lines+markers"}: Curve display.
- show_brier
Logical: Include each sample's Brier score in its legend label.
- rug
Logical: Show the distribution of individual probabilities.
- rug_size
Numeric
(0, Inf): Rug tick height in pixels.- rug_opacity
Numeric
[0, 1]: Rug tick opacity.- point_size
Numeric
(0, Inf): Calibration marker diameter in pixels.- line_width
Numeric: Curve stroke width in pixels.
- digits
Integer: Decimal places for AUC labels and tooltip values.
- diagonal
Logical: Show an independent chance diagonal.
- diagonal_color
Character: Chance-line color.
- palette
Optional Character: Group colors; unset uses the chart theme.
- legend
Logical: Show group labels and AUC summaries.
- square
Logical: Keep the plotting grid square.
- xlab
Character: Horizontal axis label.
- ylab
Character: Vertical axis label.
Details
Compiles to LineSeriesOption in src/chart/line/LineSeries.ts and
ScatterSeriesOption in src/chart/scatter/ScatterSeries.ts.
ECharts docs: https://echarts.apache.org/en/option.html#series-line
Statistical semantics
Bins are computed independently within each group. Equidistant bins partition
the unit interval. Quantile bins use linear interpolation at positions
(n - 1) * p + 1 in the sorted probabilities (R quantile type 7).
Repeated boundaries are collapsed without splitting tied scores. Intervals
include their lower boundary and exclude their upper boundary, except that
the final interval includes its upper boundary. Constant probabilities form
one bin. Empty bins are omitted; lines join the remaining bin means.
Each point is the mean predicted probability and mean observed outcome in
its bin. The Brier score is the mean squared probability error across all
complete observations, before binning. It is not an average of bin errors.
Missing pairs are removed together when na_rm is true; invalid finite
ranges and infinite values are always rejected. A group with no complete
observations is rejected. The probability rug uses those same complete
observations and sits just inside the lower plotting edge.
Both axes span zero to one. Display precision does not round plotted values.
A one-bin curve remains visible as a point even in lines-only mode.
Examples
config <- setup_CalibrationConfig(n_bins = 2L, bin_method = "equidistant")
draw(config, data = data.frame(observed = c(0, 1), probability = c(.2, .8)))