Skip to contents

Bind binary observation records to a reliability diagram. Each row contains an observed outcome (zero or one) and its predicted probability, optionally grouped by sample or model. Configuration carries column names, not data.

Usage

CalibrationConfig(
  dat_path = NULL,
  title = NULL,
  origin = NULL,
  writer = NULL,
  legend_position = "top",
  legend_placement = "outside",
  observed = "observed",
  probability = "probability",
  group = NULL,
  n_bins = 10L,
  bin_method = "quantile",
  na_rm = TRUE,
  mode = "lines+markers",
  show_brier = TRUE,
  rug = TRUE,
  rug_size = 8,
  rug_opacity = 0.3,
  point_size = 7,
  line_width = 2,
  digits = 3L,
  diagonal = TRUE,
  diagonal_color = "#888888",
  palette = NULL,
  legend = TRUE,
  square = TRUE,
  xlab = "Mean predicted probability",
  ylab = "Observed proportion"
)

Arguments

dat_path

Optional Character: Path to the data, read at draw time. The serializable alternative to passing data to draw().

title

Optional Character: Chart title.

origin

Optional Named character {"user", "default", "derived"}: Where each value came from, one entry per settable property. Absent on an authored config; written by the interface that resolved it.

writer

Optional Named character: Which interface wrote the config, as name and version. Absent on an authored config.

legend_position

Character {"top", "bottom", "left", "right", "top-left", "top-right", "bottom-left", "bottom-right"}: Legend anchor. Top/bottom anchors use horizontal rows; left/right anchors use a vertical column. Corner anchors align within the top or bottom row.

legend_placement

Character {"outside", "inside"}: Relation to the plotting area. Outside placement reserves space for the complete legend; inside placement overlays the data. Neither setting adds a missing legend.

observed

Character: Column containing binary outcomes, zero or one.

probability

Character: Column containing predicted probabilities.

group

Optional Character: Column identifying samples or models.

n_bins

Integer [1, Inf): Requested number of bins per group.

bin_method

Character {"quantile", "equidistant"}: Binning rule.

na_rm

Logical: Remove incomplete observation/probability pairs.

mode

Character {"lines", "markers", "lines+markers"}: Curve display.

show_brier

Logical: Include each sample's Brier score in its legend label.

rug

Logical: Show the distribution of individual probabilities.

rug_size

Numeric (0, Inf): Rug tick height in pixels.

rug_opacity

Numeric [0, 1]: Rug tick opacity.

point_size

Numeric (0, Inf): Calibration marker diameter in pixels.

line_width

Numeric: Curve stroke width in pixels.

digits

Integer: Decimal places for AUC labels and tooltip values.

diagonal

Logical: Show an independent chance diagonal.

diagonal_color

Character: Chance-line color.

palette

Optional Character: Group colors; unset uses the chart theme.

legend

Logical: Show group labels and AUC summaries.

square

Logical: Keep the plotting grid square.

xlab

Character: Horizontal axis label.

ylab

Character: Vertical axis label.

Value

A CalibrationConfig object.

Details

Compiles to LineSeriesOption in src/chart/line/LineSeries.ts and ScatterSeriesOption in src/chart/scatter/ScatterSeries.ts. ECharts docs: https://echarts.apache.org/en/option.html#series-line

Statistical semantics

Bins are computed independently within each group. Equidistant bins partition the unit interval. Quantile bins use linear interpolation at positions (n - 1) * p + 1 in the sorted probabilities (R quantile type 7). Repeated boundaries are collapsed without splitting tied scores. Intervals include their lower boundary and exclude their upper boundary, except that the final interval includes its upper boundary. Constant probabilities form one bin. Empty bins are omitted; lines join the remaining bin means. Each point is the mean predicted probability and mean observed outcome in its bin. The Brier score is the mean squared probability error across all complete observations, before binning. It is not an average of bin errors. Missing pairs are removed together when na_rm is true; invalid finite ranges and infinite values are always rejected. A group with no complete observations is rejected. The probability rug uses those same complete observations and sits just inside the lower plotting edge. Both axes span zero to one. Display precision does not round plotted values. A one-bin curve remains visible as a point even in lines-only mode.

Examples

config <- setup_CalibrationConfig(n_bins = 2L, bin_method = "equidistant")
draw(config, data = data.frame(observed = c(0, 1), probability = c(.2, .8)))