# `BatchClassifier`

### *class* capymoa.base.BatchClassifier[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_batch_classifier.py#L19)

Bases: [`Classifier`](capymoa.base.Classifier.md#capymoa.base.Classifier), [`Batch`](capymoa.base.Batch.md#capymoa.base.Batch), [`ABC`](https://docs.python.org/3/library/abc.html#abc.ABC)

Base class for classifiers that support mini-batches.

Supported by:

- [`capymoa.ocl.evaluation.ocl_train_eval_loop()`](capymoa.ocl.evaluation.md#capymoa.ocl.evaluation.ocl_train_eval_loop)
- [`capymoa.evaluation.prequential_evaluation()`](capymoa.evaluation.md#capymoa.evaluation.prequential_evaluation)

Evaluators that support batch classifiers will call the [`batch_train()`](#capymoa.base.BatchClassifier.batch_train)
and [`batch_predict_proba()`](#capymoa.base.BatchClassifier.batch_predict_proba) methods instead of [`train()`](#capymoa.base.BatchClassifier.train) and
[`predict_proba()`](#capymoa.base.BatchClassifier.predict_proba):

```pycon
>>> from capymoa.base import BatchClassifier
>>> from capymoa.datasets import ElectricityTiny
>>> from capymoa.evaluation import prequential_evaluation
>>>
>>> batch_size = 500
>>> class MyBatchClassifier(BatchClassifier):
...     def batch_train(self, x, y):
...         print(f"batch_train x: {x.shape} {x.dtype}")
...         print(f"batch_train y: {y.shape} {y.dtype}")
...
...     def batch_predict_proba(self, x):
...         print(f"batch_predict_proba x: {x.shape} {x.dtype}")
...         return torch.zeros((x.shape[0], self.schema.get_num_classes()))
...
>>> stream = ElectricityTiny()
>>> learner = MyBatchClassifier(stream.get_schema())
>>> _ = prequential_evaluation(
...     stream,
...     learner,
...     batch_size=batch_size,
...     max_instances=721
... )
batch_predict_proba x: torch.Size([500, 6]) torch.float32
batch_train x: torch.Size([500, 6]) torch.float32
batch_train y: torch.Size([500]) torch.int64
batch_predict_proba x: torch.Size([221, 6]) torch.float32
batch_train x: torch.Size([221, 6]) torch.float32
batch_train y: torch.Size([221]) torch.int64
```

You can manually use `itertools.batched` (python 3.12) function and
`np.stack` to collect batches of instances as a matrix:

```pycon
>>> from itertools import islice
>>> from capymoa._utils import batched # Not available in python < 3.12
>>> stream.restart() # streams are stateful, so restart it
>>> for i, batch in enumerate(batched(stream, 100)):
...     x = np.stack([instance.x for instance in batch])
...     y = np.stack([instance.y_index for instance in batch])
...     x = torch.from_numpy(x).to(learner.device, learner.x_dtype)
...     y = torch.from_numpy(y).to(learner.device, learner.y_dtype)
...     learner.batch_train(x, y)
...     break
batch_train x: torch.Size([100, 6]) torch.float32
batch_train y: torch.Size([100]) torch.int64
```

The default implementation of [`train()`](#capymoa.base.BatchClassifier.train) and [`predict()`](#capymoa.base.BatchClassifier.predict) calls the
batch variants with a batch of size 1. This is useful for parts of CapyMOA
that expect a classifier to be able to train and predict on single
instances.

```pycon
>>> instance = next(stream)
>>> learner.train(instance)
batch_train x: torch.Size([1, 6]) torch.float32
batch_train y: torch.Size([1]) torch.int64
>>> learner.predict(instance)
batch_predict_proba x: torch.Size([1, 6]) torch.float32
0
>>> learner.predict_proba(instance)
batch_predict_proba x: torch.Size([1, 6]) torch.float32
array([0., 0.])
```

#### \_\_init_\_(schema: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema), random_seed: [int](https://docs.python.org/3/builtins/functions.html#int) = 1)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_classifier.py#L29)

#### batch_predict(x: [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)) → [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_batch_classifier.py#L115)

Predict the labels for a batch of instances.

* **Parameters:**
  **x** – Batch of [`x_dtype`](#capymoa.base.BatchClassifier.x_dtype) valued feature vectors
  `(batch_size, num_features)`
* **Returns:**
  Predicted batch of [`y_dtype`](#capymoa.base.BatchClassifier.y_dtype) valued labels
  `(batch_size,)`.

#### *abstract* batch_predict_proba(x: [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)) → [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_batch_classifier.py#L105)

Predict the probabilities of the classes for a batch of instances.

* **Parameters:**
  **x** – Batch of [`x_dtype`](#capymoa.base.BatchClassifier.x_dtype) valued feature vectors
  `(batch_size, num_features)`
* **Returns:**
  Batch of [`x_dtype`](#capymoa.base.BatchClassifier.x_dtype) valued predicted probabilities
  `(batch_size, num_classes)`.

#### *abstract* batch_train(x: [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor), y: [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_batch_classifier.py#L96)

Train with a batch of instances.

* **Parameters:**
  * **x** – Batch of [`x_dtype`](#capymoa.base.BatchClassifier.x_dtype) valued feature vectors
    `(batch_size, num_features)`
  * **y** – Batch of [`y_dtype`](#capymoa.base.BatchClassifier.y_dtype) valued labels `(batch_size,)`.

#### *classmethod* from_params(schema: [Any](https://docs.python.org/3/library/typing.html#typing.Any) = None, params: [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Any](https://docs.python.org/3/library/typing.html#typing.Any)] | [None](https://docs.python.org/3/builtins/constants.html#None) = None, random_seed: [int](https://docs.python.org/3/builtins/functions.html#int) = 1) → [Any](https://docs.python.org/3/library/typing.html#typing.Any)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_learner_params.py#L170)

Construct an instance from parameters produced by `get_params`.

#### get_params() → [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Any](https://docs.python.org/3/library/typing.html#typing.Any)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_learner_params.py#L163)

Return the hyper-parameters captured from the constructor.

#### predict(instance: [Instance](capymoa.core.Instance.md#capymoa.core.Instance)) → [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_classifier.py#L56)

Predict the label of an instance.

The base implementation calls [`predict_proba()`](#capymoa.base.BatchClassifier.predict_proba) and returns the
label with the highest probability.

* **Parameters:**
  **instance** – The instance to predict the label for.
* **Returns:**
  The predicted label or `None` if the classifier is unable
  to make a prediction.

#### predict_proba(instance: [Instance](capymoa.core.Instance.md#capymoa.core.Instance)) → [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Any](https://docs.python.org/3/library/typing.html#typing.Any), ...], [dtype](https://numpy.org/doc/stable/reference/generated/numpy.dtype.html#numpy.dtype)[float64]] | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_batch_classifier.py#L134)

Calls [`batch_predict_proba()`](#capymoa.base.BatchClassifier.batch_predict_proba) with a batch of size 1.

#### train(instance: [LabeledInstance](capymoa.core.LabeledInstance.md#capymoa.core.LabeledInstance)) → [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_batch_classifier.py#L125)

Calls [`batch_train()`](#capymoa.base.BatchClassifier.batch_train) with a batch of size 1.

#### device *: [torch.device](https://docs.pytorch.org/docs/stable/tensor_attributes.html#torch.device)* *= device(type='cpu')*

Device on which the batch will be processed.

#### random_seed *: [int](https://docs.python.org/3/builtins/functions.html#int)*

The random seed for reproducibility.

When implementing a classifier ensure random number generators are seeded.

#### schema *: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)*

The schema representing the instances.

#### x_dtype *: [dtype](https://docs.pytorch.org/docs/stable/tensor_attributes.html#torch.dtype)* *= torch.float32*

Data type for the input features.

#### y_dtype *: [dtype](https://docs.pytorch.org/docs/stable/tensor_attributes.html#torch.dtype)* *= torch.int64*

Data type for the target value/labels.
