# `Nomao`

### *class* capymoa.datasets.Nomao[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_datasets.py#L367)

Bases: `_DownloadableARFF`

Nomao is a deduplication classification problem based on data aggregated
by the Nomao Labs search engine.

* Number of instances: 34,465
* Number of attributes: 118
* Number of classes: 2

The task is to determine whether two place records (containing information
such as names, phone numbers, and addresses collected from several sources)
refer to the same real-world place. The attributes measure similarity and
matching characteristics across the various fields of the records being compared.

**References:**

1. Candillier, Laurent, and Vincent Lemaire. “Design and analysis of the
   Nomao challenge active learning in the real-world.” Proceedings of the
   ALRA: Active Learning in Real-world Applications, Workshop ECML-PKDD
   (2012): [https://archive.ics.uci.edu/ml/datasets/Nomao](https://archive.ics.uci.edu/ml/datasets/Nomao).
2. “nomao.” OpenML (2015): [https://www.openml.org/d/1486](https://www.openml.org/d/1486).

#### \_\_init_\_(directory: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, auto_download: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True, file_type: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['arff', 'csv'] = 'arff')[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L76)

Setup a stream from a dataset file and optionally download it if missing.

* **Parameters:**
  * **directory** – Where downloads are stored.
    Defaults to [`capymoa.datasets.get_download_dir()`](capymoa.datasets.md#capymoa.datasets.get_download_dir).
  * **auto_download** – Download the dataset if it is missing.
  * **file_type** – Download either the `"arff"` or `"csv"` dataset asset.

#### \_\_iter_\_() → [Self](https://docs.python.org/3/library/typing.html#typing.Self)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L356)

Get an iterator over the stream.

This will NOT restart the stream if it has already been iterated over.
Please use the [`restart()`](#capymoa.datasets.Nomao.restart) method to restart the stream.

* **Yield:**
  An iterator over the stream.

#### \_\_next_\_() → \_AnyInstance[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L366)

Get the next instance in the stream.

* **Returns:**
  The next instance in the stream.

#### cli_help() → [str](https://docs.python.org/3/builtins/stdtypes.html#str)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L379)

Return a help message

#### get_moa_stream() → InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L113)

Get the MOA stream object if it exists.

#### get_schema() → [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L110)

Return the schema of the stream.

#### has_more_instances() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L104)

Return `True` if the stream have more instances to read.

#### next_instance()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L107)

Return the next instance in the stream.

* **Raises:**
  [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If the machine learning task is neither a regression
  nor a classification task.
* **Returns:**
  A labeled instances or a regression depending on the schema.

#### restart()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L116)

Restart the stream to read instances from the beginning.

#### *classmethod* to_stream(path: [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path)) → [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/datasets/_downloader.py#L96)

Convert the downloaded and unpacked dataset into a datastream.

#### moa_stream *: \_InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)*

#### schema *: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)*

#### stream *: [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)*
