# `KDD99`

### *class* capymoa.datasets.KDD99[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_datasets.py#L342)

Bases: `_DownloadableARFF`

KDD99 is a network intrusion detection problem based on a 10% stratified
subsample of the 1998 DARPA Intrusion Detection Evaluation Program data.

* Number of instances: 494,020
* Number of attributes: 41
* Number of classes: 23

The task is to distinguish between normal connections and different types of
network intrusions (attacks), grouped into four main categories: denial-of-service,
unauthorized access from a remote machine, unauthorized access to local
superuser privileges, and surveillance/probing.

**References:**

1. Stolfo, Salvatore, Wei Fan, Wenke Lee, Andreas Prodromidis, and Philip
   Chan. “KDD Cup 1999 Data.” UCI Machine Learning Repository (1999):
   [https://doi.org/10.24432/C51C7N](https://doi.org/10.24432/C51C7N).
2. “KDDCup99.” OpenML (2014): [https://www.openml.org/d/1113](https://www.openml.org/d/1113).

#### \_\_init_\_(directory: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, auto_download: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True, file_type: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['arff', 'csv'] = 'arff')[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L76)

Setup a stream from a dataset file and optionally download it if missing.

* **Parameters:**
  * **directory** – Where downloads are stored.
    Defaults to [`capymoa.datasets.get_download_dir()`](capymoa.datasets.md#capymoa.datasets.get_download_dir).
  * **auto_download** – Download the dataset if it is missing.
  * **file_type** – Download either the `"arff"` or `"csv"` dataset asset.

#### \_\_iter_\_() → [Self](https://docs.python.org/3/library/typing.html#typing.Self)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L356)

Get an iterator over the stream.

This will NOT restart the stream if it has already been iterated over.
Please use the [`restart()`](#capymoa.datasets.KDD99.restart) method to restart the stream.

* **Yield:**
  An iterator over the stream.

#### \_\_next_\_() → \_AnyInstance[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L366)

Get the next instance in the stream.

* **Returns:**
  The next instance in the stream.

#### cli_help() → [str](https://docs.python.org/3/builtins/stdtypes.html#str)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L379)

Return a help message

#### get_moa_stream() → InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L113)

Get the MOA stream object if it exists.

#### get_schema() → [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L110)

Return the schema of the stream.

#### has_more_instances() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L104)

Return `True` if the stream have more instances to read.

#### next_instance()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L107)

Return the next instance in the stream.

* **Raises:**
  [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If the machine learning task is neither a regression
  nor a classification task.
* **Returns:**
  A labeled instances or a regression depending on the schema.

#### restart()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L116)

Restart the stream to read instances from the beginning.

#### *classmethod* to_stream(path: [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path)) → [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L96)

Convert the downloaded and unpacked dataset into a datastream.

#### moa_stream *: \_InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)*

#### schema *: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)*

#### stream *: [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)*
