BaseDataDriftDetector#

class capymoa.drift.detectors.BaseDataDriftDetector[source]#

Bases: BaseDriftDetector

Base class for detectors that monitor the input data distribution.

These detectors have REQUIRES_FIT set to True: they need a reference distribution. Provide it with fit(), or construct with auto_fit_samples so add_element() collects it.

Subclasses set the class variable IS_UNIVARIATE and implement:

  • _fit() – store or preprocess the reference data.

  • _test() – run the test (on one feature when univariate, on all features when multivariate).

  • get_params() – return detector hyper-parameters.

The base class handles the sliding test window, feature-wise looping for univariate tests, and detection bookkeeping. For univariate tests that produce p-values, it also applies Bonferroni correction [1] across features. Univariate tests that compare a distance or score to a fixed threshold instead (no p-value) supply their own per-feature decisions directly; correction has no effect on those.

By default the reference window is fixed after fit() (or after auto-fit) [2] [3]. Call fit() again along the stream to use a sliding reference.

__init__(
window_size: int,
alpha: float = 0.05,
correction: Literal['bonferroni', 'none'] = 'bonferroni',
auto_fit_samples: int | None = None,
)[source]#

Create a data drift detector.

Parameters:
  • window_size – Number of observations to collect before running a comparison against the reference. The detector returns no result (and detected_change is False) until the window is full.

  • alpha – Significance level. For p-value tests, drift is declared when the (corrected) p-value falls below alpha.

  • correction – Multiple-testing correction for univariate tests that produce p-values, across features. "bonferroni" (default) divides alpha by the number of features; "none" uses alpha directly. Ignored for multivariate tests, and for univariate tests that compare a distance or score to a fixed threshold instead of a p-value (those supply their own per-feature decisions, uncorrected).

  • auto_fit_samples – If set, the first auto_fit_samples observations are used as the reference (auto-fit mode). No explicit fit() call is needed; add_element() collects the reference. If None (default), fit() must be called before add_element() or compare().

Raises:

ValueError – If window_size is not a positive integer, alpha is not in (0, 1], correction is unknown, or auto_fit_samples is not a positive integer when set.

abstract _fit(X: ndarray) → None[source]#

Store or preprocess the reference data.

Called by fit() after validation. X is guaranteed to be a 2-D np.ndarray with shape (n_samples, n_features).

At minimum, implementations should set self._X_ref = X.

abstract _test(
X_ref: ndarray,
X_test: ndarray,
) → DataDriftResult[source]#

Run the underlying statistical test or distance measure.

When IS_UNIVARIATE = True this receives one feature at a time: both arrays are 1-D with shapes (n_ref,) and (n_test,). The subclass should set statistic and p_value (if available) on the returned result.

If p_value is set, the base class applies the configured correction and ignores the per-feature is_drift flag. Example:

def _test(self, x_ref, x_test):
    stat, p = scipy.stats.ks_2samp(x_ref, x_test)
    return DataDriftResult(
        is_drift=False, statistic=stat, p_value=p
    )

If p_value is left None (distance- or score-based tests), the base class uses the per-feature is_drift flag directly and correction has no effect – the subclass is responsible for its own per-feature decision (e.g. comparing a distance to a threshold).

When IS_UNIVARIATE = False this receives all features: both arrays are 2-D with shapes (n_ref, n_features) and (n_test, n_features). The subclass is responsible for the full DataDriftResult including is_drift.

add_element(element: float | ndarray) → None[source]#

Add one observation and check for drift.

The observation is appended to a sliding window of size window_size. Once full, the detector compares the window against the reference on every call.

If the detector is not yet fitted and auto_fit_samples is set, those initial observations build the reference. No comparison happens until then. If it is not fitted and auto-fit is not enabled, this raises RuntimeError.

Parameters:

element – A single observation – a scalar for univariate data, or a 1-D array for multivariate data.

Raises:
  • RuntimeError – If the detector is not fitted and auto-fit mode is not enabled. Call fit() first, or construct with auto_fit_samples.

  • ValueError – If the observation has a different number of features than the reference.

compare(
X_test: ndarray,
) → DataDriftResult[source]#

One-shot batch comparison against the reference.

Unlike add_element(), this does not update the internal window or detection history. It is useful for offline evaluation. auto_fit_samples does not apply here; call fit() first.

Parameters:

X_test – Test data with the same number of features as the reference.

Raises:
Returns:

Comparison result.

detected_change() → bool[source]#

Is the detector currently detecting a concept drift?

detected_warning() → bool[source]#

Is the detector currently warning of an upcoming concept drift?

fit(
X: ndarray | Any,
feature_names: Sequence[str] | None = None,
) → None[source]#

Set the reference distribution.

This detector has REQUIRES_FIT set to True. Call fit() before add_element() or compare(), unless auto_fit_samples was set so add_element() can collect the reference. Calling fit() again replaces the reference (a sliding reference window).

Parameters:
  • X – Reference data, shape (n_samples,) or (n_samples, n_features). May also be a pandas DataFrame, in which case column names are extracted automatically.

  • feature_names – Optional names for each feature. If provided, must have length equal to the number of features. Overrides column names extracted from a DataFrame. When set, per-feature dicts in DataDriftResult use these names as keys instead of integer indices.

Raises:

ValueError – If X is empty or feature_names length does not match the number of features.

classmethod from_params(
schema: Any = None,
params: dict[str, Any] | None = None,
random_seed: int = 1,
) → Any[source]#

Construct an instance from parameters produced by get_params.

abstract get_params() → dict[str, Any][source]#

Return the hyper-parameters of this detector.

reset(clean_history: bool = False) → None[source]#

Reset the detector state.

Parameters:

clean_history – If True, also clear the reference data and detection history. If False (default), only the sliding window and current result are cleared; the reference and detection indices are preserved.

IS_UNIVARIATE: bool = True#

If True the test is applied to each feature separately and results are combined. If False the test runs on the joint distribution.

REQUIRES_FIT: bool = True#

Data drift detectors need a reference distribution. Call fit() or set auto_fit_samples so add_element() can collect it.

property X_ref: ndarray | None#

The reference data set with fit.

property alpha: float#

Significance level for drift decisions.

property auto_fit_samples: int | None#

Number of samples for auto-fit, or None if explicit fit.

property correction: str#

Multiple-testing correction across features.

property feature_names: list[str] | None#

Feature names passed to fit(), or None.

property is_fitted: bool#

Whether the reference distribution has been set.

True after fit(), or after add_element() has collected auto_fit_samples observations.

property n_features: int | None#

Number of features in the reference data, or None before fit.

property result: DataDriftResult | None#

Most recent comparison result, or None during warm-up.

property window_size: int#

Size of the sliding test window.