# `KolmogorovSmirnov`

### *class* capymoa.drift.detectors.KolmogorovSmirnov[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/ks.py#L11)

Bases: [`BaseDataDriftDetector`](capymoa.drift.detectors.BaseDataDriftDetector.md#capymoa.drift.detectors.BaseDataDriftDetector)

Kolmogorov-Smirnov Two-Sample Test

Univariate statistical test that compares the empirical CDFs of the
reference and test samples. Large differences in the CDF indicate
that the two samples come from different distributions.

Applied per feature, with Bonferroni correction by default.

## Example:

```pycon
>>> import numpy as np
>>> from capymoa.drift.detectors import KolmogorovSmirnov
>>> rng = np.random.default_rng(42)
>>> detector = KolmogorovSmirnov(window_size=50)
>>> detector.fit(rng.normal(0, 1, size=(200, 2)))
>>> for x in rng.normal(3, 1, size=(50, 2)):
...     detector.add_element(x)
>>> detector.detected_change()
True
```

## Reference:

Massey Jr, F. J. “The Kolmogorov-Smirnov test for goodness of fit.”
Journal of the American Statistical Association 46.253 (1951): 68-78.

#### \_\_init_\_(window_size: [int](https://docs.python.org/3/builtins/functions.html#int), alpha: [float](https://docs.python.org/3/builtins/functions.html#float) = 0.05, correction: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['bonferroni', 'none'] = 'bonferroni', auto_fit_samples: [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, alternative: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['two-sided', 'less', 'greater'] = 'two-sided', method: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['auto', 'exact', 'approx', 'asymp'] = 'auto')[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/ks.py#L43)

Create a Kolmogorov-Smirnov data drift detector.

* **Parameters:**
  * **window_size** – Number of observations in the sliding window.
  * **alpha** – Significance level for drift decisions.
  * **correction** – Multiple-testing correction across features.
  * **auto_fit_samples** – Number of initial samples for auto-fit.
  * **alternative** – Alternative hypothesis for the KS test.
    `"two-sided"` (default), `"less"`, or `"greater"`.
  * **method** – Method for computing the p-value. See
    [`scipy.stats.ks_2samp()`](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.ks_2samp.html#scipy.stats.ks_2samp) for details.

#### \_fit(X: [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray)) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/ks.py#L67)

Store or preprocess the reference data.

Called by [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) after validation. *X* is guaranteed to be a
2-D `np.ndarray` with shape `(n_samples, n_features)`.

At minimum, implementations should set `self._X_ref = X`.

#### \_test(x_ref: [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray), x_test: [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray)) → [DataDriftResult](capymoa.drift.detectors.DataDriftResult.md#capymoa.drift.detectors.DataDriftResult)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/ks.py#L70)

Run the underlying statistical test or distance measure.

When `IS_UNIVARIATE = True` this receives **one feature** at a
time: both arrays are 1-D with shapes `(n_ref,)` and
`(n_test,)`. The subclass should set `statistic` and
`p_value` (if available) on the returned result.

If `p_value` is set, the base class applies the configured
`correction` and ignores the per-feature `is_drift` flag.
Example:

```default
def _test(self, x_ref, x_test):
    stat, p = scipy.stats.ks_2samp(x_ref, x_test)
    return DataDriftResult(
        is_drift=False, statistic=stat, p_value=p
    )
```

If `p_value` is left `None` (distance- or score-based tests),
the base class uses the per-feature `is_drift` flag directly and
`correction` has no effect – the subclass is responsible for its
own per-feature decision (e.g. comparing a distance to a
threshold).

When `IS_UNIVARIATE = False` this receives **all features**:
both arrays are 2-D with shapes `(n_ref, n_features)` and
`(n_test, n_features)`. The subclass is responsible for
the full `DataDriftResult` including `is_drift`.

#### add_element(element: [float](https://docs.python.org/3/builtins/functions.html#float) | [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray)) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/base.py#L300)

Add one observation and check for drift.

The observation is appended to a sliding window of size
[`window_size`](#capymoa.drift.detectors.KolmogorovSmirnov.window_size). Once full, the detector compares the window
against the reference on every call.

If the detector is not yet fitted and [`auto_fit_samples`](#capymoa.drift.detectors.KolmogorovSmirnov.auto_fit_samples)
is set, those initial observations build the reference.
No comparison happens until then. If it is not fitted and
auto-fit is not enabled, this raises [`RuntimeError`](https://docs.python.org/3/builtins/exceptions.html#RuntimeError).

* **Parameters:**
  **element** – A single observation – a scalar for univariate
  data, or a 1-D array for multivariate data.
* **Raises:**
  * [**RuntimeError**](https://docs.python.org/3/builtins/exceptions.html#RuntimeError) – If the detector is not fitted and
    auto-fit mode is not enabled. Call [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) first, or
    construct with `auto_fit_samples`.
  * [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If the observation has a different number of
    features than the reference.

#### compare(X_test: [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray)) → [DataDriftResult](capymoa.drift.detectors.DataDriftResult.md#capymoa.drift.detectors.DataDriftResult)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/base.py#L354)

One-shot batch comparison against the reference.

Unlike [`add_element()`](#capymoa.drift.detectors.KolmogorovSmirnov.add_element), this does not update the internal
window or detection history. It is useful for offline evaluation.
`auto_fit_samples` does not apply here; call [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) first.

* **Parameters:**
  **X_test** – Test data with the same number of features as the
  reference.
* **Raises:**
  * [**RuntimeError**](https://docs.python.org/3/builtins/exceptions.html#RuntimeError) – If [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) has not been called.
  * [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If *X_test* has a different number of features
    than the reference.
* **Returns:**
  Comparison result.

#### detected_change() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/base_detector.py#L52)

Is the detector currently detecting a concept drift?

#### detected_warning() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/base_detector.py#L56)

Is the detector currently warning of an upcoming concept drift?

#### fit(X: [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray) | [Any](https://docs.python.org/3/library/typing.html#typing.Any), feature_names: [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)] | [None](https://docs.python.org/3/builtins/constants.html#None) = None) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/base.py#L254)

Set the reference distribution.

This detector has [`REQUIRES_FIT`](#capymoa.drift.detectors.KolmogorovSmirnov.REQUIRES_FIT) set to `True`. Call
[`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) before [`add_element()`](#capymoa.drift.detectors.KolmogorovSmirnov.add_element) or [`compare()`](#capymoa.drift.detectors.KolmogorovSmirnov.compare), unless
`auto_fit_samples` was set so [`add_element()`](#capymoa.drift.detectors.KolmogorovSmirnov.add_element) can collect
the reference. Calling [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) again replaces the reference
(a sliding reference window).

* **Parameters:**
  * **X** – Reference data, shape `(n_samples,)` or
    `(n_samples, n_features)`. May also be a pandas DataFrame,
    in which case column names are extracted automatically.
  * **feature_names** – Optional names for each feature. If
    provided, must have length equal to the number of features.
    Overrides column names extracted from a DataFrame.
    When set, per-feature dicts in [`DataDriftResult`](capymoa.drift.detectors.DataDriftResult.md#capymoa.drift.detectors.DataDriftResult) use
    these names as keys instead of integer indices.
* **Raises:**
  [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If *X* is empty or *feature_names* length
  does not match the number of features.

#### *classmethod* from_params(schema: [Any](https://docs.python.org/3/library/typing.html#typing.Any) = None, params: [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Any](https://docs.python.org/3/library/typing.html#typing.Any)] | [None](https://docs.python.org/3/builtins/constants.html#None) = None, random_seed: [int](https://docs.python.org/3/builtins/functions.html#int) = 1) → [Any](https://docs.python.org/3/library/typing.html#typing.Any)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_learner_params.py#L170)

Construct an instance from parameters produced by `get_params`.

#### get_params() → [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Any](https://docs.python.org/3/library/typing.html#typing.Any)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/ks.py#L79)

Return the hyper-parameters of this detector.

#### reset(clean_history: [bool](https://docs.python.org/3/builtins/functions.html#bool) = False) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/drift/detectors/data_drift/base.py#L382)

Reset the detector state.

* **Parameters:**
  **clean_history** – If `True`, also clear the reference data
  and detection history. If `False` (default), only the
  sliding window and current result are cleared; the reference
  and detection indices are preserved.

#### IS_UNIVARIATE *: [bool](https://docs.python.org/3/builtins/functions.html#bool)* *= True*

If `True` the test is applied to each feature separately and
results are combined. If `False` the test runs on the joint
distribution.

#### REQUIRES_FIT *: [bool](https://docs.python.org/3/builtins/functions.html#bool)* *= True*

Data drift detectors need a reference distribution. Call
[`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit) or set `auto_fit_samples` so [`add_element()`](#capymoa.drift.detectors.KolmogorovSmirnov.add_element)
can collect it.

#### *property* X_ref *: [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray) | [None](https://docs.python.org/3/builtins/constants.html#None)*

The reference data set with `fit`.

#### *property* alpha *: [float](https://docs.python.org/3/builtins/functions.html#float)*

Significance level for drift decisions.

#### *property* auto_fit_samples *: [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None)*

Number of samples for auto-fit, or `None` if explicit fit.

#### *property* correction *: [str](https://docs.python.org/3/builtins/stdtypes.html#str)*

Multiple-testing correction across features.

#### *property* feature_names *: [list](https://docs.python.org/3/builtins/stdtypes.html#list)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)] | [None](https://docs.python.org/3/builtins/constants.html#None)*

Feature names passed to [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit), or `None`.

#### *property* is_fitted *: [bool](https://docs.python.org/3/builtins/functions.html#bool)*

Whether the reference distribution has been set.

`True` after [`fit()`](#capymoa.drift.detectors.KolmogorovSmirnov.fit), or after [`add_element()`](#capymoa.drift.detectors.KolmogorovSmirnov.add_element) has
collected [`auto_fit_samples`](#capymoa.drift.detectors.KolmogorovSmirnov.auto_fit_samples) observations.

#### *property* n_features *: [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None)*

Number of features in the reference data, or `None` before fit.

#### *property* result *: [DataDriftResult](capymoa.drift.detectors.DataDriftResult.md#capymoa.drift.detectors.DataDriftResult) | [None](https://docs.python.org/3/builtins/constants.html#None)*

Most recent comparison result, or `None` during warm-up.

#### *property* window_size *: [int](https://docs.python.org/3/builtins/functions.html#int)*

Size of the sliding test window.
