# `datasets`

CapyMOA comes with some datasets ‘out of the box’. Simply import the dataset
and start using it, the data will be downloaded automatically if it is not
already present in the download directory. You can configure where the datasets
are downloaded to by setting an environment variable (See [`capymoa.env`](capymoa.env.md#module-capymoa.env))

```pycon
>>> from capymoa.datasets import ElectricityTiny
>>> stream = ElectricityTiny()
>>> stream.next_instance().x
array([0.      , 0.056443, 0.439155, 0.003467, 0.422915, 0.414912])
```

## Classes

| [`KDD99`](capymoa.datasets.KDD99.md#capymoa.datasets.KDD99)                     | KDD99 is a network intrusion detection problem based on a 10% stratified subsample of the 1998 DARPA Intrusion Detection Evaluation Program data.                            |
|---------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| [`Airlines`](capymoa.datasets.Airlines.md#capymoa.datasets.Airlines)               | Airlines dataset inspired in the regression dataset from Elena Ikonomovska.                                                                                                  |
| [`Bike`](capymoa.datasets.Bike.md#capymoa.datasets.Bike)                       | Bike is a regression dataset for the amount of bike share information.                                                                                                       |
| [`CovtFD`](capymoa.datasets.CovtFD.md#capymoa.datasets.CovtFD)                   | CovtFD is an adaptation from the classic [`Covtype`](capymoa.datasets.Covtype.md#capymoa.datasets.Covtype) classification problem with added feature drifts. |
| [`Covtype`](capymoa.datasets.Covtype.md#capymoa.datasets.Covtype)                 | The classic covertype (/covtype) classification problem                                                                                                                      |
| [`CovtypeNorm`](capymoa.datasets.CovtypeNorm.md#capymoa.datasets.CovtypeNorm)         | A normalized version of the classic [`Covtype`](capymoa.datasets.Covtype.md#capymoa.datasets.Covtype) classification problem.                                |
| [`CovtypeTiny`](capymoa.datasets.CovtypeTiny.md#capymoa.datasets.CovtypeTiny)         | A truncated version of the classic [`Covtype`](capymoa.datasets.Covtype.md#capymoa.datasets.Covtype) classification problem.                                 |
| [`Electricity`](capymoa.datasets.Electricity.md#capymoa.datasets.Electricity)         | Electricity is a classification problem based on the Australian New South Wales Electricity Market.                                                                          |
| [`ElectricityTiny`](capymoa.datasets.ElectricityTiny.md#capymoa.datasets.ElectricityTiny) | A truncated version of the Electricity dataset with 1000 instances.                                                                                                          |
| [`Fried`](capymoa.datasets.Fried.md#capymoa.datasets.Fried)                     | Fried is a regression problem based on the Friedman dataset.                                                                                                                 |
| [`FriedTiny`](capymoa.datasets.FriedTiny.md#capymoa.datasets.FriedTiny)             | A truncated version of the Friedman regression problem with 1000 instances.                                                                                                  |
| [`Hyper100k`](capymoa.datasets.Hyper100k.md#capymoa.datasets.Hyper100k)             | Hyper100k is a classification problem based on the moving hyperplane generator.                                                                                              |
| [`Nomao`](capymoa.datasets.Nomao.md#capymoa.datasets.Nomao)                     | Nomao is a deduplication classification problem based on data aggregated by the Nomao Labs search engine.                                                                    |
| [`PokerHand`](capymoa.datasets.PokerHand.md#capymoa.datasets.PokerHand)             | PokerHand is a classification problem where each instance is an example of a hand consisting of five playing cards drawn from a standard deck of 52.                         |
| [`RBFm_100k`](capymoa.datasets.RBFm_100k.md#capymoa.datasets.RBFm_100k)             | RBFm_100k is a synthetic classification problem based on the Radial Basis Function generator.                                                                                |
| [`RTG_2abrupt`](capymoa.datasets.RTG_2abrupt.md#capymoa.datasets.RTG_2abrupt)         | RTG_2abrupt is a synthetic classification problem based on the Random Tree generator with 2 abrupt drifts.                                                                   |
| [`Sensor`](capymoa.datasets.Sensor.md#capymoa.datasets.Sensor)                   | Sensor stream is a classification problem based on indoor sensor data.                                                                                                       |
| [`Spambase`](capymoa.datasets.Spambase.md#capymoa.datasets.Spambase)               | Spambase is a classification problem based on the classic UCI Spambase email dataset from Hewlett-Packard Labs.                                                              |

## Functions

| [`download_unpacked`](#capymoa.datasets.download_unpacked)   | Download and unpack an archived/compressed file.                                                                          |
|----------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------|
| [`get_download_dir`](#capymoa.datasets.get_download_dir)    | Get a directory where datasets should be downloaded to.                                                                   |
| [`load_openml_dataset`](#capymoa.datasets.load_openml_dataset) | Load any OpenML dataset by numeric id as a [`Stream`](capymoa.stream.Stream.md#capymoa.stream.Stream). |

### capymoa.datasets.download_unpacked(url: [str](https://docs.python.org/3/builtins/stdtypes.html#str), downloads: [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path) | [str](https://docs.python.org/3/builtins/stdtypes.html#str)) → [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_utils.py#L47)

Download and unpack an archived/compressed file.

* `https://example.com/mydir.tar.gz` -> `/downloads/mydir/`
* `https://example.com/mydir.zip` -> `/downloads/mydir/`
* `https://example.com/myfile.gz` -> `/downloads/myfile`

* **Parameters:**
  * **url** – URL with a valid filename.
  * **downloads** – Base directory to download to.
* **Raises:**
  * [**FileNotFoundError**](https://docs.python.org/3/builtins/exceptions.html#FileNotFoundError) – If `downloads` does not exist.
  * [**NotADirectoryError**](https://docs.python.org/3/builtins/exceptions.html#NotADirectoryError) – If `downloads` is not a directory.
  * [**FileExistsError**](https://docs.python.org/3/builtins/exceptions.html#FileExistsError) – If the unpacked directory/file already exists.
* **Returns:**
  The unpacked directory/file.

### capymoa.datasets.get_download_dir(download_dir: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [None](https://docs.python.org/3/builtins/constants.html#None) = None) → [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_utils.py#L89)

Get a directory where datasets should be downloaded to.

The download directory is determined by the following steps:

1. If the `download_dir` parameter is provided, use that.
2. If the `CAPYMOA_DATASETS_DIR` environment variable is set, use that.
3. Otherwise, use the default download directory: `./data`.

* **Parameters:**
  **download_dir** – Override the download directory.
* **Returns:**
  The download directory.

### capymoa.datasets.load_openml_dataset(openml_id: [int](https://docs.python.org/3/builtins/functions.html#int), directory: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, auto_download: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True) → [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_openml.py#L62)

Load any OpenML dataset by numeric id as a [`Stream`](capymoa.stream.Stream.md#capymoa.stream.Stream).

```pycon
>>> from capymoa.datasets import load_openml_dataset
>>> stream = load_openml_dataset(61)  # iris
>>> stream.next_instance()
LabeledInstance(
    Schema(iris),
    x=[5.1 3.5 1.4 0.2],
    y_index=0,
    y_label='Iris-setosa'
)
```

* **Parameters:**
  * **openml_id** – The numeric OpenML dataset id, e.g. `1169` for
    [https://www.openml.org/d/1169](https://www.openml.org/d/1169).
  * **directory** – Where downloads are stored. Defaults to
    [`capymoa.datasets.get_download_dir()`](#capymoa.datasets.get_download_dir).
  * **auto_download** – Fetch metadata / download the dataset if missing.
* **Raises:**
  * [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If the dataset is not stored as ARFF on OpenML, or its
    target attribute is missing, ambiguous, or not found in the ARFF file.
  * [**ConnectionError**](https://docs.python.org/3/builtins/exceptions.html#ConnectionError) – If the OpenML API cannot be reached.
  * [**FileNotFoundError**](https://docs.python.org/3/builtins/exceptions.html#FileNotFoundError) – If the metadata or dataset is missing locally
    and `auto_download` is False.
