# `CovtFD`

### *class* capymoa.datasets.CovtFD[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_datasets.py#L44)

Bases: `_DownloadableARFF`

CovtFD is an adaptation from the classic [`Covtype`](capymoa.datasets.Covtype.md#capymoa.datasets.Covtype) classification
problem with added feature drifts.

* Number of instances: 581,011 (30m^2 cells)
* Number of attributes: 104 (10 continuous, 44 categorical, 50 dummy)
* Number of classes: 7 (forest cover types)

Given 30x30-meter cells obtained from the US Resource Information System
(RIS). The dataset includes 10 continuous and 44 categorical features, which
we augmented by adding 50 dummy continuous features drawn from a Normal
probability distribution with μ = 0 and σ = 1. Only the continuous features
were randomly swapped with 10 (out of the fifty) dummy features to simulate
drifts. We added such synthetic drift twice, one at instance 193, 669 and
another at 387, 338.

**References:**

1. Gomes, Heitor Murilo, Rodrigo Fernandes de Mello, Bernhard Pfahringer,
   and Albert Bifet. “Feature scoring using tree-based ensembles for
   evolving data streams.” In 2019 IEEE International Conference on
   Big Data (Big Data), pp. 761-769. IEEE, 2019.
2. Blackard,Jock. (1998). Covertype. UCI Machine Learning Repository.
   [https://doi.org/10.24432/C50K5N](https://doi.org/10.24432/C50K5N).
3. [https://archive.ics.uci.edu/ml/datasets/Covertype](https://archive.ics.uci.edu/ml/datasets/Covertype)

**See Also:**

* [`Covtype`](capymoa.datasets.Covtype.md#capymoa.datasets.Covtype) - The classic covertype dataset
* [`CovtypeNorm`](capymoa.datasets.CovtypeNorm.md#capymoa.datasets.CovtypeNorm) - A normalized version of the classic covertype dataset
* [`CovtypeTiny`](capymoa.datasets.CovtypeTiny.md#capymoa.datasets.CovtypeTiny) - A truncated version of the classic covertype dataset

#### \_\_init_\_(directory: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, auto_download: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True, file_type: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['arff', 'csv'] = 'arff')[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L76)

Setup a stream from a dataset file and optionally download it if missing.

* **Parameters:**
  * **directory** – Where downloads are stored.
    Defaults to [`capymoa.datasets.get_download_dir()`](capymoa.datasets.md#capymoa.datasets.get_download_dir).
  * **auto_download** – Download the dataset if it is missing.
  * **file_type** – Download either the `"arff"` or `"csv"` dataset asset.

#### \_\_iter_\_() → [Self](https://docs.python.org/3/library/typing.html#typing.Self)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L356)

Get an iterator over the stream.

This will NOT restart the stream if it has already been iterated over.
Please use the [`restart()`](#capymoa.datasets.CovtFD.restart) method to restart the stream.

* **Yield:**
  An iterator over the stream.

#### \_\_next_\_() → \_AnyInstance[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L366)

Get the next instance in the stream.

* **Returns:**
  The next instance in the stream.

#### cli_help() → [str](https://docs.python.org/3/builtins/stdtypes.html#str)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L379)

Return a help message

#### get_moa_stream() → InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L113)

Get the MOA stream object if it exists.

#### get_schema() → [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L110)

Return the schema of the stream.

#### has_more_instances() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L104)

Return `True` if the stream have more instances to read.

#### next_instance()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L107)

Return the next instance in the stream.

* **Raises:**
  [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If the machine learning task is neither a regression
  nor a classification task.
* **Returns:**
  A labeled instances or a regression depending on the schema.

#### restart()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L116)

Restart the stream to read instances from the beginning.

#### *classmethod* to_stream(path: [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path)) → [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L96)

Convert the downloaded and unpacked dataset into a datastream.

#### moa_stream *: \_InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)*

#### schema *: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)*

#### stream *: [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)*
