# `stream`

Data stream representations and related utilities.

## Modules

| [`drift`](capymoa.stream.drift.md#module-capymoa.stream.drift)                 | Simulate concept drift in datastreams.   |
|----------------------------------------------------------------------------------------------------|------------------------------------------|
| [`generator`](capymoa.stream.generator.md#module-capymoa.stream.generator)         | Generate artificial data streams.        |
| [`preprocessing`](capymoa.stream.preprocessing.md#module-capymoa.stream.preprocessing) | Data stream preprocessing and pipelines. |

## Classes

| [`ARFFStream`](capymoa.stream.ARFFStream.md#capymoa.stream.ARFFStream)   | A datastream originating from an ARFF file.           |
|-----------------------------------------------------------------------------------------|-------------------------------------------------------|
| [`CSVStream`](capymoa.stream.CSVStream.md#capymoa.stream.CSVStream)     | Create a CapyMOA datastream from a CSV file.          |
| [`MOAStream`](capymoa.stream.MOAStream.md#capymoa.stream.MOAStream)     | A datastream that can be learnt instance by instance. |
| [`NumpyStream`](capymoa.stream.NumpyStream.md#capymoa.stream.NumpyStream) | A datastream originating from a numpy array.          |
| [`Schema`](capymoa.stream.Schema.md#capymoa.stream.Schema)           | Schema describes the structure of a stream.           |
| [`Stream`](capymoa.stream.Stream.md#capymoa.stream.Stream)           | A datastream that can be learnt instance by instance. |
| [`TorchStream`](capymoa.stream.TorchStream.md#capymoa.stream.TorchStream) | A stream adapter for PyTorch datasets.                |

## Functions

| [`stream_from_file`](#capymoa.stream.stream_from_file)   | Create a datastream from a csv or arff file.   |
|---------------------------------------------------------------------|------------------------------------------------|

### capymoa.stream.stream_from_file(path_to_csv_or_arff: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [Path](https://docs.python.org/3/library/pathlib.html#pathlib.Path), dataset_name: [str](https://docs.python.org/3/builtins/stdtypes.html#str) = 'NoName', class_index: [int](https://docs.python.org/3/builtins/functions.html#int) = -1, target_type: [Literal](https://docs.python.org/3/library/typing.html#typing.Literal)['numeric', 'categorical'] | [None](https://docs.python.org/3/builtins/constants.html#None) = None) → [Stream](capymoa.stream.Stream.md#capymoa.stream.Stream)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream_from_file.py#L10)

Create a datastream from a csv or arff file.

```pycon
>>> from capymoa.stream import stream_from_file
>>> stream = stream_from_file(
...     "data/electricity_tiny.csv",
...     dataset_name="Electricity",
...     target_type="categorical"
... )
>>> stream.next_instance()
LabeledInstance(
    Schema(Electricity),
    x=[0.    0.056 0.439 0.003 0.423 0.415],
    y_index=1,
    y_label='1'
)
>>> stream.next_instance().x
array([0.021277, 0.051699, 0.415055, 0.003467, 0.422915, 0.414912])
```

#### SEE ALSO
* [`CSVStream`](capymoa.stream.CSVStream.md#capymoa.stream.CSVStream)
* [`ARFFStream`](capymoa.stream.ARFFStream.md#capymoa.stream.ARFFStream)
* [`NumpyStream`](capymoa.stream.NumpyStream.md#capymoa.stream.NumpyStream)
* [`TorchStream`](capymoa.stream.TorchStream.md#capymoa.stream.TorchStream)
* [`Stream`](capymoa.stream.Stream.md#capymoa.stream.Stream)

**CSV File Considerations:**

* Assumes a **header row** with attribute names.
* Supports only **numeric** and **categorical** attributes.
* String columns are automatically converted to **categorical** attributes. Convert
  them to numeric beforehand if needed.
* The whole CSV file is **read into memory**. For very large files, use
  [`CSVStream`](capymoa.stream.CSVStream.md#capymoa.stream.CSVStream) directly for streaming from disk.
* **Default Target:** The last column (class_index=-1) is assumed to be the target
  variable.
* **Missing Values:** Represented by `?`.

**ARFF File Considerations:**

Reads the Attribute-Relation File Format ([ARFF](https://waikato.github.io/weka-wiki/formats_and_processing/arff_stable/))
commonly used in MOA and WEKA.

* Supports only `NUMERIC` and `NOMINAL` attributes.
* **Missing Values:** Represented by `?`.
* **Default Target:** The last attribute is assumed to be the target variable.

* **Parameters:**
  * **path_to_csv_or_arff** – A file path to a CSV or ARFF file.
  * **dataset_name** – A descriptive name given to the dataset, defaults to “NoName”
  * **class_index** – The index of the column containing the class label. By default,
    the algorithm assumes that the class label is located in the column specified by
    this index. However, if the class label is located in a different column, you
    can specify its index using this parameter.
  * **target_type** – When working with a CSV file, this parameter allows the user to
    specify the target values in the data to be interpreted as categorical or
    numeric. Defaults to None to detect automatically.
