# `CSVStream`

### *class* capymoa.stream.CSVStream[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_csv_stream.py#L13)

Bases: [`Stream`](capymoa.stream.Stream.md#capymoa.stream.Stream)[`_AnyInstance`]

Create a CapyMOA datastream from a CSV file.

* The CSV file **must** have a header row with feature names.
* Integers or strings can specify nominal features.
* `?` represent missing values.
* CSV is read line by line, so it can handle large files.

When ‘categories’ are provided for the target attribute, then the stream returns
[`LabeledInstance`](capymoa.core.LabeledInstance.md#capymoa.core.LabeledInstance) objects.

```pycon
>>> from io import StringIO
>>> from capymoa.stream import CSVStream
>>> csv_content = '''feature1,feature2,target
... 1,A,yes
... 2,B,no
... 3,0,0
... 5,1,1
... ?,?,?
... '''
>>> csv_file = StringIO(csv_content)
>>> stream = CSVStream(
...     file=csv_file,
...     target="target",
...     categories={"target": ["yes", "no"], "feature2": ["A", "B"]},
...     name="TestStream"
... )
>>> for instance in stream:
...     print(instance.x, instance.y_index, instance.y_label)
[1. 0.] 0 yes
[2. 1.] 1 no
[3. 0.] 0 yes
[5. 1.] 1 no
[nan nan] -1 None
```

When no categories are provided for the target attribute, then the stream returns
[`RegressionInstance`](capymoa.core.RegressionInstance.md#capymoa.core.RegressionInstance) objects.

```pycon
>>> csv_content = '''target,feature1,feature2
... 0.0,A,1
... 0.5,B,2
... 1.5,0,3
... 2.0,1,4
... ?,?,?
... '''
>>> csv_file = StringIO(csv_content)
>>> stream = CSVStream(
...     file=csv_file,
...     target="target",
...     categories={"feature1": ["A", "B"]},
...     name="TestStream"
... )
>>> for instance in stream:
...     print(instance.x, instance.y_value)
[0. 1.] 0.0
[1. 2.] 0.5
[0. 3.] 1.5
[1. 4.] 2.0
[nan nan] nan
```

#### \_\_init_\_(file: [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path) | [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [TextIO](https://docs.python.org/3/library/typing.html#typing.TextIO), target: [str](https://docs.python.org/3/builtins/stdtypes.html#str), categories: [Mapping](https://docs.python.org/3/library/collections.abc.html#collections.abc.Mapping)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)]] | [None](https://docs.python.org/3/builtins/constants.html#None) = None, name: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, length: [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None) = None) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_csv_stream.py#L75)

Create a CSV stream.

* **Parameters:**
  * **file** – A path to a CSV file or an open file-like object.
  * **target** – The name of the target attribute.
  * **categories** – A mapping from attribute names to their categorical values.
  * **name** – An optional name for the stream. If not provided, the filename is
    used.
  * **length** – An optional length of the stream (number of instances). If provided,
    this enables the [`Sized`](https://docs.python.org/3/library/typing.html#typing.Sized) interface.

#### \_\_iter_\_() → [Self](https://docs.python.org/3/library/typing.html#typing.Self)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L356)

Get an iterator over the stream.

This will NOT restart the stream if it has already been iterated over.
Please use the [`restart()`](#capymoa.stream.CSVStream.restart) method to restart the stream.

* **Yield:**
  An iterator over the stream.

#### \_\_next_\_() → \_AnyInstance[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L366)

Get the next instance in the stream.

* **Returns:**
  The next instance in the stream.

#### cli_help() → [str](https://docs.python.org/3/builtins/stdtypes.html#str)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L379)

Return a help message

#### get_moa_stream() → InstanceStream | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L402)

Get the MOA stream object if it exists.

#### get_schema() → [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_csv_stream.py#L135)

Return the schema of the stream.

#### has_more_instances() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_csv_stream.py#L139)

Return `True` if the stream have more instances to read.

#### next_instance() → \_AnyInstance[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_csv_stream.py#L143)

Return the next instance in the stream.

* **Raises:**
  [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – If the machine learning task is neither a regression
  nor a classification task.
* **Returns:**
  A labeled instances or a regression depending on the schema.

#### restart() → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_csv_stream.py#L151)

Restart the stream to read instances from the beginning.
