# `Schema`

### *class* capymoa.stream.Schema[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L50)

Bases: [`object`](https://docs.python.org/3/builtins/functions.html#object)

Schema describes the structure of a stream.

It contains the attribute names, datatype, and the possible values for nominal attributes.
The schema is crucial for a learner to know how to interpret instances correctly.

When working with datasets built into CapyMOA (see [`capymoa.datasets`](capymoa.datasets.md#module-capymoa.datasets))
and ARFF files, the schema is automatically created. However, in some cases
you might want to create a schema manually. This can be done using the
[`from_custom()`](#capymoa.stream.Schema.from_custom) method.

#### \_\_init_\_(moa_header: InstancesHeader)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L62)

Construct a schema by wrapping a `InstancesHeader`.

To create a schema without an `InstancesHeader` use
[`from_custom()`](#capymoa.stream.Schema.from_custom) method.

* **Parameters:**
  **moa_header** – A Java MOA header object.

#### describe_difference(other: [Schema](#capymoa.stream.Schema)) → [list](https://docs.python.org/3/builtins/stdtypes.html#list)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L186)

Describe how this schema differs from `other`.

* **Parameters:**
  **other** – The schema to compare against.
* **Returns:**
  A list of human-readable differences, empty if the two schemas
  are compatible.

```pycon
>>> from capymoa.datasets import ElectricityTiny, FriedTiny
>>> ElectricityTiny().get_schema().describe_difference(
...     ElectricityTiny().get_schema()
... )
[]
>>> for line in ElectricityTiny().get_schema().describe_difference(
...     FriedTiny().get_schema()
... ):
...     print(line)
number of attributes: 6 != 10
numeric attribute names differ
task: classification != regression
```

#### *static* from_custom(features: [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)], target: [str](https://docs.python.org/3/builtins/stdtypes.html#str), categories: [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)]] | [None](https://docs.python.org/3/builtins/constants.html#None) = None, name: [str](https://docs.python.org/3/builtins/stdtypes.html#str) = 'unnamed')[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L277)

Create a CapyMOA Schema that defines each attribute in the stream.

The following example shows how to use this method to create a classification
schema:

```pycon
>>> from capymoa.stream import Schema
>>> schema = Schema.from_custom(
...     features=["f1", "f2", "class"],
...     target="class",
...     categories={"class": ["yes", "no"], "f1": ["low", "medium", "high"]},
...     name="classification-example"
... )
>>> print(schema)
@relation classification-example

@attribute f1 {low,medium,high}
@attribute f2 numeric
@attribute class {yes,no}

@data
>>> print(schema.is_classification())
True
```

The following example shows how to use this method to create a regression
schema:

```pycon
>>> schema = Schema.from_custom(
...     features=["f1", "f2", "target"],
...     target="target",
...     categories={"f1": ["A", "B", "C"]},
...     name="regression-example"
... )
>>> print(schema)
@relation regression-example

@attribute f1 {A,B,C}
@attribute f2 numeric
@attribute target numeric

@data
>>> print(schema.is_regression())
True
```

* **Parameters:**
  * **features** – A list of feature names.
  * **target** – The name of the target attribute. Must be in features as well.
  * **categories** – A dictionary mapping feature names to their possible values.
    When the target attribute is included in this dictionary the task is
    considered classification.
  * **name** – The name of the dataset.
* **Returns:**
  A CapyMOA Schema object.

#### get_index_for_label(y: [str](https://docs.python.org/3/builtins/stdtypes.html#str))[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L113)

Return the index for the class label y.

#### get_label_indexes() → [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[int](https://docs.python.org/3/builtins/functions.html#int)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L101)

Return the possible indexes for the class label.

#### get_label_values() → [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L96)

Return the possible values for the class label.

#### get_moa_header() → InstancesHeader[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L118)

Get the JAVA MOA header. Useful for advanced users.

This is needed for advanced operations that are not supported by the
Python wrappers (yet).

#### get_nominal_attributes() → [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)]][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L148)

Return a dict of nominal attributes.

#### get_num_attributes() → [int](https://docs.python.org/3/builtins/functions.html#int)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L126)

Return the number of attributes excluding the target attribute.

#### get_num_classes() → [int](https://docs.python.org/3/builtins/functions.html#int)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L167)

Return the number of possible classes. If regression, returns 1.

#### get_num_nominal_attributes() → [int](https://docs.python.org/3/builtins/functions.html#int)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L130)

Return the number of nominal attributes.

#### get_num_numeric_attributes() → [int](https://docs.python.org/3/builtins/functions.html#int)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L139)

Return the number of numeric attributes.

#### get_numeric_attributes() → [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[str](https://docs.python.org/3/builtins/stdtypes.html#str)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L158)

Return a list of numeric attribute names.

#### get_value_for_index(y_index: [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None)) → [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L106)

Return the value for the class label index y_index.

#### is_classification() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L177)

Return True if the problem is a classification problem.

#### is_compatible_with(other: [Schema](#capymoa.stream.Schema)) → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L231)

Return True if instances of `other` can be consumed where this schema is expected.

Compares attribute structure and target, ignoring the dataset name. Use
[`describe_difference()`](#capymoa.stream.Schema.describe_difference) to find out *what* differs.

* **Parameters:**
  **other** – The schema to compare against.

```pycon
>>> from capymoa.datasets import ElectricityTiny, FriedTiny
>>> ElectricityTiny().get_schema().is_compatible_with(
...     ElectricityTiny().get_schema()
... )
True
>>> ElectricityTiny().get_schema().is_compatible_with(
...     FriedTiny().get_schema()
... )
False
```

#### is_regression() → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L173)

Return True if the problem is a regression problem.

#### is_y_index_in_range(y_index: [int](https://docs.python.org/3/builtins/functions.html#int)) → [bool](https://docs.python.org/3/builtins/functions.html#bool)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/stream/_stream.py#L181)

Return True if the y_index is in the range of the class label indexes.

#### *property* dataset_name *: [str](https://docs.python.org/3/builtins/stdtypes.html#str)*

Returns the name of the dataset.

#### *property* shape *: [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[int](https://docs.python.org/3/builtins/functions.html#int)]*

The shape of the input `x` instances.

Usually [`capymoa.core.Instance.x`](capymoa.core.Instance.md#capymoa.core.Instance.x) is a vector but some learners
need to know the shape of the input. For example, a CNN needs to know the height
and width of an image.
