ABCD#

class capymoa.drift.detectors.ABCD[source]#

Bases: BaseDriftDetector

Adaptive Bernstein Change Detector (ABCD).

ABCD is a drift detector for multivariate data streams. It fits an encoder-decoder model of the incoming feature vectors, monitors how well that model reconstructs them, and signals a change when the reconstruction error shifts by more than a Bernstein-bound-based test allows. Because it watches the input distribution rather than a single value, it detects drift that changes the features, not drift that only changes the labels. The reconstruction model is selected with model_id: "pca" (default) and "kpca" rely only on scikit-learn, while "ae" uses a PyTorch autoencoder and therefore needs the torch extra.

Example:#

>>> from capymoa.drift.detectors import ABCD
>>> from capymoa.stream.drift import AbruptDrift, DriftStream
>>> from capymoa.stream.generator import RandomRBFGenerator
>>>
>>> def rbf(seed):
...     return RandomRBFGenerator(
...         model_random_seed=seed,
...         instance_random_seed=seed,
...         number_of_attributes=6,
...         number_of_centroids=20,
...     )
...
>>> # A stream whose input distribution changes at instance 2000.
>>> stream = DriftStream(stream=[rbf(1), AbruptDrift(position=2000), rbf(99)])
>>> detector = ABCD(model_id="pca", maximum_absolute_value=0.2)
>>>
>>> i = 0
>>> while stream.has_more_instances() and i < 4000:
...     instance = stream.next_instance()
...     i += 1
...     detector.add_element(instance.x)
...     if detector.detected_change():
...         print(f"Change detected at index: {i}")
Change detected at index: 2324

Reference:#

Heyden, M., Fouché, E., Arzamasov, V., Fenn, T., Kalinke, F., & Böhm, K. (2024). Adaptive Bernstein change detector for high-dimensional data streams. Data Mining and Knowledge Discovery, 38(3), 1334-1363. Springer.

__init__(
delta_drift: float = 0.002,
delta_warn: float = 0.01,
model_id: str = 'pca',
split_type: str = 'ed',
encoding_factor: float = 0.5,
update_epochs: int = 50,
num_splits: int = 20,
max_size: int = np.inf,
subspace_threshold: float = 2.5,
n_min: int = 100,
maximum_absolute_value: float = 1.0,
bonferroni: bool = False,
)[source]#
Parameters:
  • delta_drift – The desired confidence level at which a drift is detected

  • delta_warn – The desired confidence level at which a warning is detected

  • model_id – Which encoder-decoder to reconstruct instances with: "pca", "kpca" or "ae". "pca" is the default because it needs only scikit-learn, whereas "ae" is backed by PyTorch, which CapyMOA installs only as the torch extra.

  • update_epochs – The number of epochs to train the AE after a change occurred

  • split_type – Investigation of different split types

  • subspace_threshold – Called tau in the paper

  • bonferroni – Use bonferroni correction to account for multiple testing?

  • encoding_factor – The relative size of the bottleneck

  • maximum_absolute_value – The maximum absolute value that one can expect (e.g. 1.0 for normalized data). Smaller values can increase false alarms but speed up change detection

  • num_splits – The number of time point to evaluate

  • max_size – The maximum number of instances kept in the adaptive window; older instances beyond this size are discarded. Defaults to no limit.

  • n_min – The number of initial instances collected to pre-train the reconstruction model before change monitoring begins

add_element(
element: float | ndarray | Instance,
) → None[source]#

Add the new element and also perform change detection :param element: The new observation :return:

detected_change() → bool[source]#

Is the detector currently detecting a concept drift?

detected_warning() → bool[source]#

Is the detector currently warning of an upcoming concept drift?

classmethod from_params(
schema: Any = None,
params: dict[str, Any] | None = None,
random_seed: int = 1,
) → Any[source]#

Construct an instance from parameters produced by get_params.

get_dims_p_values() → ndarray[source]#
get_drift_dims() → ndarray[source]#
get_params() → dict[str, Any][source]#

Get the hyper-parameters of the drift detector.

get_severity()[source]#
loss()[source]#
pre_train(data)[source]#
reset(clean_history: bool = False) → None[source]#

Reset the detector state.

Clears the adaptive window, the reconstruction model, and all bookkeeping so that replaying the same input reproduces the as-constructed flag trace. Hyper-parameters (including model_id) are preserved; the reconstruction model is re-trained lazily from the next n_min elements, exactly as on construction.

Parameters:

clean_history – Whether to reset detection history, defaults to False

REQUIRES_FIT: bool = False#

If True, this detector needs a reference distribution before it can detect change. Concept drift detectors leave this as False and only use add_element().