Spambase#
- class capymoa.datasets.Spambase[source]#
Bases:
_DownloadableARFFSpambase is a classification problem based on the classic UCI Spambase email dataset from Hewlett-Packard Labs.
Number of instances: 4,601
Number of attributes: 57
Number of classes: 2
The task is to predict whether a given email is spam. Spam emails were collected from postmaster and individual submissions, while non-spam emails came from filed work and personal correspondence. Attributes consist of word and character frequency measures as well as measures of the length of consecutive sequences of capital letters.
References:
Hopkins, Mark, Erik Reeber, George Forman, and Jaap Suermondt. “Spambase.” UCI Machine Learning Repository (1999): https://archive.ics.uci.edu/dataset/94/spambase, https://doi.org/10.24432/C53G6X.
“spambase.” OpenML (2014): https://www.openml.org/d/44.
- __init__(
- directory: str | Path | None = None,
- auto_download: bool = True,
- file_type: Literal['arff', 'csv'] = 'arff',
Setup a stream from a dataset file and optionally download it if missing.
- Parameters:
directory – Where downloads are stored. Defaults to
capymoa.datasets.get_download_dir().auto_download – Download the dataset if it is missing.
file_type – Download either the
"arff"or"csv"dataset asset.
- __iter__() Self[source]#
Get an iterator over the stream.
This will NOT restart the stream if it has already been iterated over. Please use the
restart()method to restart the stream.- Yield:
An iterator over the stream.
- __next__() _AnyInstance[source]#
Get the next instance in the stream.
- Returns:
The next instance in the stream.
- next_instance()[source]#
Return the next instance in the stream.
- Raises:
ValueError – If the machine learning task is neither a regression nor a classification task.
- Returns:
A labeled instances or a regression depending on the schema.