Nomao#
- class capymoa.datasets.Nomao[source]#
Bases:
_DownloadableARFFNomao is a deduplication classification problem based on data aggregated by the Nomao Labs search engine.
Number of instances: 34,465
Number of attributes: 118
Number of classes: 2
The task is to determine whether two place records (containing information such as names, phone numbers, and addresses collected from several sources) refer to the same real-world place. The attributes measure similarity and matching characteristics across the various fields of the records being compared.
References:
Candillier, Laurent, and Vincent Lemaire. “Design and analysis of the Nomao challenge active learning in the real-world.” Proceedings of the ALRA: Active Learning in Real-world Applications, Workshop ECML-PKDD (2012): https://archive.ics.uci.edu/ml/datasets/Nomao.
“nomao.” OpenML (2015): https://www.openml.org/d/1486.
- __init__(
- directory: str | Path | None = None,
- auto_download: bool = True,
- file_type: Literal['arff', 'csv'] = 'arff',
Setup a stream from a dataset file and optionally download it if missing.
- Parameters:
directory – Where downloads are stored. Defaults to
capymoa.datasets.get_download_dir().auto_download – Download the dataset if it is missing.
file_type – Download either the
"arff"or"csv"dataset asset.
- __iter__() Self[source]#
Get an iterator over the stream.
This will NOT restart the stream if it has already been iterated over. Please use the
restart()method to restart the stream.- Yield:
An iterator over the stream.
- __next__() _AnyInstance[source]#
Get the next instance in the stream.
- Returns:
The next instance in the stream.
- next_instance()[source]#
Return the next instance in the stream.
- Raises:
ValueError – If the machine learning task is neither a regression nor a classification task.
- Returns:
A labeled instances or a regression depending on the schema.