# `data`

Utilities for continual learning when using PyTorch datasets.

## Functions

| [`class_incremental_schedule`](#capymoa.ocl.util.data.class_incremental_schedule)   | Returns a class schedule for class incremental learning.             |
|-------------------------------------------------------------------------------|----------------------------------------------------------------------|
| [`class_incremental_split`](#capymoa.ocl.util.data.class_incremental_split)      | Divide a dataset into multiple tasks for class incremental learning. |
| [`class_schedule_to_task_mask`](#capymoa.ocl.util.data.class_schedule_to_task_mask)  | Convert a class schedule to a list of boolean masks.                 |
| [`get_targets`](#capymoa.ocl.util.data.get_targets)                  | Return the targets of a dataset as a 1D tensor.                      |
| [`group_indicies`](#capymoa.ocl.util.data.group_indicies)               | Group the indices of a dataset according to a class schedule.        |
| [`partition_by_schedule`](#capymoa.ocl.util.data.partition_by_schedule)        | Divide a dataset into multiple datasets based on a class schedule.   |

### capymoa.ocl.util.data.class_incremental_schedule(num_classes: [int](https://docs.python.org/3/builtins/functions.html#int), num_tasks: [int](https://docs.python.org/3/builtins/functions.html#int), shuffle: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True, generator: [Generator](https://docs.pytorch.org/docs/stable/generated/torch.Generator.html#torch.Generator) = torch.default_generator) → [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[set](https://docs.python.org/3/builtins/stdtypes.html#set)[[int](https://docs.python.org/3/builtins/functions.html#int)]][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/ocl/util/data.py#L158)

Returns a class schedule for class incremental learning.

```pycon
>>> class_incremental_schedule(9, 3, shuffle=False)
[{0, 1, 2}, {3, 4, 5}, {8, 6, 7}]
```

```pycon
>>> class_incremental_schedule(9, 3, generator=torch.Generator().manual_seed(0))
[{8, 0, 2}, {1, 3, 7}, {4, 5, 6}]
```

* **Parameters:**
  * **num_classes** – The number of classes in the dataset.
  * **num_tasks** – The number of tasks to divide the classes into.
  * **shuffle** – When False, the classes occur in numerical order of their
    labels. When True, the classes are shuffled.
  * **generator** – The random number generator used for shuffling, defaults
    to torch.default_generator
* **Returns:**
  A list of lists of classes for each task.

### capymoa.ocl.util.data.class_incremental_split(dataset: [Dataset](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.Dataset)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor), [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)]], num_tasks: [int](https://docs.python.org/3/builtins/functions.html#int), shuffle_tasks: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True, generator: [Generator](https://docs.pytorch.org/docs/stable/generated/torch.Generator.html#torch.Generator) = torch.default_generator) → [tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[Dataset](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.Dataset)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor), [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)]]], [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[set](https://docs.python.org/3/builtins/stdtypes.html#set)[[int](https://docs.python.org/3/builtins/functions.html#int)]]][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/ocl/util/data.py#L118)

Divide a dataset into multiple tasks for class incremental learning.

```pycon
>>> from torch.utils.data import TensorDataset
>>> x = torch.tensor([[1, 2], [3, 4], [5, 6], [7, 8]])
>>> y = torch.tensor([0, 1, 2, 3])
>>> dataset = TensorDataset(x, y)
>>> tasks, schedule = class_incremental_split(dataset, 2, shuffle_tasks=False)
>>> schedule
[{0, 1}, {2, 3}]
>>> tasks[0][0]
(tensor([1, 2]), tensor(0))
>>> tasks[1][0]
(tensor([5, 6]), tensor(2))
```

* **Parameters:**
  * **dataset** – The dataset to divide.
  * **num_tasks** – The number of tasks to divide the dataset into.
  * **shuffle_tasks** – When False, the classes occur in numerical order of their
    labels. When True, the classes are shuffled.
  * **generator** – The random number generator used for shuffling, defaults
    to torch.default_generator
* **Returns:**
  A tuple containing the list of tasks and the class schedule.

### capymoa.ocl.util.data.class_schedule_to_task_mask(class_schedule: [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[set](https://docs.python.org/3/builtins/stdtypes.html#set)[[int](https://docs.python.org/3/builtins/functions.html#int)]], num_classes: [int](https://docs.python.org/3/builtins/functions.html#int)) → BoolTensor[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/ocl/util/data.py#L196)

Convert a class schedule to a list of boolean masks.

This is useful when implementing multi-headed neural networks for task
incremental learning.

```pycon
>>> class_schedule_to_task_mask([{0, 1}, {2, 3}], 4)
tensor([[ True,  True, False, False],
        [False, False,  True,  True]])
```

* **Parameters:**
  * **num_classes** – The total number of classes.
  * **class_schedule** – A sequence of sets containing class indices defining
    task order and composition.
* **Returns:**
  A boolean mask of shape (num_tasks, num_classes)

### capymoa.ocl.util.data.get_targets(dataset: [Dataset](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.Dataset)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor), [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)]]) → LongTensor[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/ocl/util/data.py#L29)

Return the targets of a dataset as a 1D tensor.

* If the dataset has a targets attribute, it is used.
* Otherwise, the targets are extracted from the dataset by iterating over it.

* **Parameters:**
  **dataset** – The dataset to get the targets from.
* **Returns:**
  A 1D tensor containing the targets of the dataset.

### capymoa.ocl.util.data.group_indicies(labels: LongTensor, schedule: [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[int](https://docs.python.org/3/builtins/functions.html#int)]], shuffle: [bool](https://docs.python.org/3/builtins/functions.html#bool) = False, rng: [Generator](https://docs.pytorch.org/docs/stable/generated/torch.Generator.html#torch.Generator) = torch.default_generator) → [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[LongTensor][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/ocl/util/data.py#L55)

Group the indices of a dataset according to a class schedule.

```pycon
>>> labels = torch.tensor([0, 0, 1, 1, 2, 2])
>>> schedule = [{0, 1}, {2}]
>>> group_indicies(labels, schedule)
[tensor([0, 1, 2, 3]), tensor([4, 5])]
```

* **Parameters:**
  * **labels** – A 1D tensor containing the class labels.
  * **schedule** – A sequence of sets containing class indices defining
    task order and composition.
  * **shuffle** – If True, the indices in each group are shuffled.
  * **rng** – The random number generator used for shuffling, defaults
    to torch.default_generator
* **Returns:**
  A list of tensors containing the indices for each group.

### capymoa.ocl.util.data.partition_by_schedule(dataset: [Dataset](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.Dataset)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor), [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)]], class_schedule: [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[set](https://docs.python.org/3/builtins/stdtypes.html#set)[[int](https://docs.python.org/3/builtins/functions.html#int)]], shuffle: [bool](https://docs.python.org/3/builtins/functions.html#bool) = False, rng: [Generator](https://docs.pytorch.org/docs/stable/generated/torch.Generator.html#torch.Generator) = torch.default_generator) → [Sequence](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[Dataset](https://docs.pytorch.org/docs/stable/data.html#torch.utils.data.Dataset)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor), [Tensor](https://docs.pytorch.org/docs/stable/tensors.html#torch.Tensor)]]][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/ocl/util/data.py#L90)

Divide a dataset into multiple datasets based on a class schedule.

In class incremental learning, a task is a dataset containing a subset of
the classes in the original dataset. This function divides a dataset into
multiple tasks, each containing a subset of the classes.

* **Parameters:**
  * **dataset** – The dataset to divide.
  * **class_schedule** – A sequence of sets containing class indices defining
    task order and composition.
  * **shuffle** – If True, the samples in each task are shuffled.
  * **rng** – The random number generator used for shuffling, defaults
    to torch.default_generator
* **Returns:**
  A list of datasets, each corresponding to a task.
