BaseDataDriftDetector#
- class capymoa.drift.detectors.BaseDataDriftDetector[source]#
Bases:
BaseDriftDetectorBase class for detectors that monitor the input data distribution.
These detectors have
REQUIRES_FITset toTrue: they need a reference distribution. Provide it withfit(), or construct withauto_fit_samplessoadd_element()collects it.Subclasses set the class variable
IS_UNIVARIATEand implement:_fit()– store or preprocess the reference data._test()– run the test (on one feature when univariate, on all features when multivariate).get_params()– return detector hyper-parameters.
The base class handles the sliding test window, feature-wise looping for univariate tests, and detection bookkeeping. For univariate tests that produce p-values, it also applies Bonferroni correction [1] across features. Univariate tests that compare a distance or score to a fixed threshold instead (no p-value) supply their own per-feature decisions directly;
correctionhas no effect on those.By default the reference window is fixed after
fit()(or after auto-fit) [2] [3]. Callfit()again along the stream to use a sliding reference.- __init__(
- window_size: int,
- alpha: float = 0.05,
- correction: Literal['bonferroni', 'none'] = 'bonferroni',
- auto_fit_samples: int | None = None,
Create a data drift detector.
- Parameters:
window_size – Number of observations to collect before running a comparison against the reference. The detector returns no result (and
detected_changeisFalse) until the window is full.alpha – Significance level. For p-value tests, drift is declared when the (corrected) p-value falls below
alpha.correction – Multiple-testing correction for univariate tests that produce p-values, across features.
"bonferroni"(default) dividesalphaby the number of features;"none"usesalphadirectly. Ignored for multivariate tests, and for univariate tests that compare a distance or score to a fixedthresholdinstead of a p-value (those supply their own per-feature decisions, uncorrected).auto_fit_samples – If set, the first auto_fit_samples observations are used as the reference (auto-fit mode). No explicit
fit()call is needed;add_element()collects the reference. IfNone(default),fit()must be called beforeadd_element()orcompare().
- Raises:
ValueError – If window_size is not a positive integer, alpha is not in
(0, 1], correction is unknown, or auto_fit_samples is not a positive integer when set.
- abstract _fit(X: ndarray) None[source]#
Store or preprocess the reference data.
Called by
fit()after validation. X is guaranteed to be a 2-Dnp.ndarraywith shape(n_samples, n_features).At minimum, implementations should set
self._X_ref = X.
- abstract _test( ) DataDriftResult[source]#
Run the underlying statistical test or distance measure.
When
IS_UNIVARIATE = Truethis receives one feature at a time: both arrays are 1-D with shapes(n_ref,)and(n_test,). The subclass should setstatisticandp_value(if available) on the returned result.If
p_valueis set, the base class applies the configuredcorrectionand ignores the per-featureis_driftflag. Example:def _test(self, x_ref, x_test): stat, p = scipy.stats.ks_2samp(x_ref, x_test) return DataDriftResult( is_drift=False, statistic=stat, p_value=p )
If
p_valueis leftNone(distance- or score-based tests), the base class uses the per-featureis_driftflag directly andcorrectionhas no effect – the subclass is responsible for its own per-feature decision (e.g. comparing a distance to a threshold).When
IS_UNIVARIATE = Falsethis receives all features: both arrays are 2-D with shapes(n_ref, n_features)and(n_test, n_features). The subclass is responsible for the fullDataDriftResultincludingis_drift.
- add_element(element: float | ndarray) None[source]#
Add one observation and check for drift.
The observation is appended to a sliding window of size
window_size. Once full, the detector compares the window against the reference on every call.If the detector is not yet fitted and
auto_fit_samplesis set, those initial observations build the reference. No comparison happens until then. If it is not fitted and auto-fit is not enabled, this raisesRuntimeError.- Parameters:
element – A single observation – a scalar for univariate data, or a 1-D array for multivariate data.
- Raises:
RuntimeError – If the detector is not fitted and auto-fit mode is not enabled. Call
fit()first, or construct withauto_fit_samples.ValueError – If the observation has a different number of features than the reference.
- compare(
- X_test: ndarray,
One-shot batch comparison against the reference.
Unlike
add_element(), this does not update the internal window or detection history. It is useful for offline evaluation.auto_fit_samplesdoes not apply here; callfit()first.- Parameters:
X_test – Test data with the same number of features as the reference.
- Raises:
RuntimeError – If
fit()has not been called.ValueError – If X_test has a different number of features than the reference.
- Returns:
Comparison result.
- fit( ) None[source]#
Set the reference distribution.
This detector has
REQUIRES_FITset toTrue. Callfit()beforeadd_element()orcompare(), unlessauto_fit_sampleswas set soadd_element()can collect the reference. Callingfit()again replaces the reference (a sliding reference window).- Parameters:
X – Reference data, shape
(n_samples,)or(n_samples, n_features). May also be a pandas DataFrame, in which case column names are extracted automatically.feature_names – Optional names for each feature. If provided, must have length equal to the number of features. Overrides column names extracted from a DataFrame. When set, per-feature dicts in
DataDriftResultuse these names as keys instead of integer indices.
- Raises:
ValueError – If X is empty or feature_names length does not match the number of features.
- classmethod from_params( ) Any[source]#
Construct an instance from parameters produced by
get_params.
- reset(clean_history: bool = False) None[source]#
Reset the detector state.
- Parameters:
clean_history – If
True, also clear the reference data and detection history. IfFalse(default), only the sliding window and current result are cleared; the reference and detection indices are preserved.
- IS_UNIVARIATE: bool = True#
If
Truethe test is applied to each feature separately and results are combined. IfFalsethe test runs on the joint distribution.
- REQUIRES_FIT: bool = True#
Data drift detectors need a reference distribution. Call
fit()or setauto_fit_samplessoadd_element()can collect it.
- property is_fitted: bool#
Whether the reference distribution has been set.
Trueafterfit(), or afteradd_element()has collectedauto_fit_samplesobservations.
- property result: DataDriftResult | None#
Most recent comparison result, or
Noneduring warm-up.