datasets#

CapyMOA comes with some datasets ‘out of the box’. Simply import the dataset and start using it, the data will be downloaded automatically if it is not already present in the download directory. You can configure where the datasets are downloaded to by setting an environment variable (See capymoa.env)

>>> from capymoa.datasets import ElectricityTiny
>>> stream = ElectricityTiny()
>>> stream.next_instance().x
array([0.      , 0.056443, 0.439155, 0.003467, 0.422915, 0.414912])

Classes#

KDD99

KDD99 is a network intrusion detection problem based on a 10% stratified subsample of the 1998 DARPA Intrusion Detection Evaluation Program data.

Airlines

Airlines dataset inspired in the regression dataset from Elena Ikonomovska.

Bike

Bike is a regression dataset for the amount of bike share information.

CovtFD

CovtFD is an adaptation from the classic Covtype classification problem with added feature drifts.

Covtype

The classic covertype (/covtype) classification problem

CovtypeNorm

A normalized version of the classic Covtype classification problem.

CovtypeTiny

A truncated version of the classic Covtype classification problem.

Electricity

Electricity is a classification problem based on the Australian New South Wales Electricity Market.

ElectricityTiny

A truncated version of the Electricity dataset with 1000 instances.

Fried

Fried is a regression problem based on the Friedman dataset.

FriedTiny

A truncated version of the Friedman regression problem with 1000 instances.

Hyper100k

Hyper100k is a classification problem based on the moving hyperplane generator.

Nomao

Nomao is a deduplication classification problem based on data aggregated by the Nomao Labs search engine.

PokerHand

PokerHand is a classification problem where each instance is an example of a hand consisting of five playing cards drawn from a standard deck of 52.

RBFm_100k

RBFm_100k is a synthetic classification problem based on the Radial Basis Function generator.

RTG_2abrupt

RTG_2abrupt is a synthetic classification problem based on the Random Tree generator with 2 abrupt drifts.

Sensor

Sensor stream is a classification problem based on indoor sensor data.

Spambase

Spambase is a classification problem based on the classic UCI Spambase email dataset from Hewlett-Packard Labs.

Functions#

download_unpacked

Download and unpack an archived/compressed file.

get_download_dir

Get a directory where datasets should be downloaded to.

load_openml_dataset

Load any OpenML dataset by numeric id as a Stream.

capymoa.datasets.download_unpacked(url: str, downloads: Path | str) → Path[source]#

Download and unpack an archived/compressed file.

  • https://example.com/mydir.tar.gz -> /downloads/mydir/

  • https://example.com/mydir.zip -> /downloads/mydir/

  • https://example.com/myfile.gz -> /downloads/myfile

Parameters:
  • url – URL with a valid filename.

  • downloads – Base directory to download to.

Raises:
Returns:

The unpacked directory/file.

capymoa.datasets.get_download_dir(download_dir: str | None = None) → Path[source]#

Get a directory where datasets should be downloaded to.

The download directory is determined by the following steps:

  1. If the download_dir parameter is provided, use that.

  2. If the CAPYMOA_DATASETS_DIR environment variable is set, use that.

  3. Otherwise, use the default download directory: ./data.

Parameters:

download_dir – Override the download directory.

Returns:

The download directory.

capymoa.datasets.load_openml_dataset(
openml_id: int,
directory: str | Path | None = None,
auto_download: bool = True,
) → Stream[source]#

Load any OpenML dataset by numeric id as a Stream.

>>> from capymoa.datasets import load_openml_dataset
>>> stream = load_openml_dataset(61)  # iris
>>> stream.next_instance()
LabeledInstance(
    Schema(iris),
    x=[5.1 3.5 1.4 0.2],
    y_index=0,
    y_label='Iris-setosa'
)
Parameters:
Raises:
  • ValueError – If the dataset is not stored as ARFF on OpenML, or its target attribute is missing, ambiguous, or not found in the ARFF file.

  • ConnectionError – If the OpenML API cannot be reached.

  • FileNotFoundError – If the metadata or dataset is missing locally and auto_download is False.