# `EFDT`

### *class* capymoa.classifier.EFDT[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/classifier/_efdt.py#L14)

Bases: [`MOAClassifier`](capymoa.base.MOAClassifier.md#capymoa.base.MOAClassifier)

Extremely Fast Decision Tree.

Extremely Fast Decision Tree (EFDT) <sup>[1](#id2)</sup> is a decision tree classifier. Also
referred to as the Hoeffding AnyTime Tree (HATT) classifier. In practice,
despite the name, EFDTs are typically slower than a vanilla Hoeffding Tree
to process data. The speed differences come from the mechanism of split re-
evaluation present in EFDT. Nonetheless, EFDT has theoretical properties
that ensure it converges faster than the vanilla Hoeffding Tree to the
structure that would be created by a batch decision tree model (such as
Classification and Regression Trees - CART). Keep in mind that such
propositions hold when processing a stationary data stream. When dealing
with non-stationary data, EFDT is somewhat robust to concept drifts as it
continually revisits and updates its internal decision tree structure.
Still, in such cases, the Hoeffding Adaptive Tree might be a better option,
as it was specifically designed to handle non-stationarity.

```pycon
>>> from capymoa.classifier import EFDT
>>> from capymoa.datasets import ElectricityTiny
>>> from capymoa.evaluation import prequential_evaluation
>>>
>>> stream = ElectricityTiny()
>>> classifier = EFDT(stream.get_schema())
>>> results = prequential_evaluation(stream, classifier, max_instances=1000)
>>> print(f"{results['cumulative'].accuracy():.1f}")
84.4
```

* <a id='id2'>**[1]**</a> [Extremely fast decision tree. Manapragada, Chaitanya, G. I. Webb, M. Salehi. ACM SIGKDD, pp. 1953-1962, 2018.](https://dl.acm.org/doi/abs/10.1145/3219819.3220005)

#### \_\_init_\_(schema: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema) | [None](https://docs.python.org/3/builtins/constants.html#None) = None, random_seed: [int](https://docs.python.org/3/builtins/functions.html#int) = 0, grace_period: [int](https://docs.python.org/3/builtins/functions.html#int) = 200, min_samples_reevaluate: [int](https://docs.python.org/3/builtins/functions.html#int) = 200, split_criterion: [str](https://docs.python.org/3/builtins/stdtypes.html#str) | [SplitCriterion](capymoa.core.moa.splitcriteria.SplitCriterion.md#capymoa.core.moa.splitcriteria.SplitCriterion) = 'InfoGainSplitCriterion', confidence: [float](https://docs.python.org/3/builtins/functions.html#float) = 1e-3, tie_threshold: [float](https://docs.python.org/3/builtins/functions.html#float) = 0.05, leaf_prediction: [str](https://docs.python.org/3/builtins/stdtypes.html#str) = 'NaiveBayesAdaptive', nb_threshold: [int](https://docs.python.org/3/builtins/functions.html#int) = 0, numeric_attribute_observer: [str](https://docs.python.org/3/builtins/stdtypes.html#str) = 'GaussianNumericAttributeClassObserver', binary_split: [bool](https://docs.python.org/3/builtins/functions.html#bool) = False, max_byte_size: [float](https://docs.python.org/3/builtins/functions.html#float) = 33554433, memory_estimate_period: [int](https://docs.python.org/3/builtins/functions.html#int) = 1000000, stop_mem_management: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True, remove_poor_attrs: [bool](https://docs.python.org/3/builtins/functions.html#bool) = False, disable_prepruning: [bool](https://docs.python.org/3/builtins/functions.html#bool) = True)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/classifier/_efdt.py#L46)

Construct an Extremely Fast Decision Tree (EFDT) Classifier

* **Parameters:**
  * **schema** – The schema of the stream.
  * **random_seed** – The random seed passed to the MOA learner.
  * **grace_period** – Number of instances a leaf should observe between split attempts.
  * **min_samples_reevaluate** – Number of instances a node should observe before re-evaluating the best split.
  * **split_criterion** – Split criterion to use. Defaults to InfoGainSplitCriterion.
  * **confidence** – Significance level to calculate the Hoeffding bound. The significance level is given by
    1 - delta. Values closer to zero imply longer split decision delays.
  * **tie_threshold** – Threshold below which a split will be forced to break ties.
  * **leaf_prediction** – Prediction mechanism used at the leaves
    (“MajorityClass” or 0, “NaiveBayes” or 1, “NaiveBayesAdaptive” or 2).
  * **nb_threshold** – Number of instances a leaf should observe before allowing Naive Bayes.
  * **numeric_attribute_observer** – The Splitter or Attribute Observer (AO) used to monitor the class statistics
    of numeric features and perform splits.
  * **binary_split** – If True, only allow binary splits.
  * **max_byte_size** – The max size of the tree, in bytes.
  * **memory_estimate_period** – Interval (number of processed instances) between memory consumption checks.
  * **stop_mem_management** – If True, stop growing as soon as memory limit is hit.
  * **remove_poor_attrs** – If True, disable poor attributes to reduce memory usage.
  * **disable_prepruning** – If True, disable merit-based tree pre-pruning.

#### cli_help()[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_classifier.py#L111)

#### *classmethod* from_params(schema: [Any](https://docs.python.org/3/library/typing.html#typing.Any) = None, params: [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Any](https://docs.python.org/3/library/typing.html#typing.Any)] | [None](https://docs.python.org/3/builtins/constants.html#None) = None, random_seed: [int](https://docs.python.org/3/builtins/functions.html#int) = 1) → [Any](https://docs.python.org/3/library/typing.html#typing.Any)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_learner_params.py#L170)

Construct an instance from parameters produced by `get_params`.

#### get_params() → [dict](https://docs.python.org/3/builtins/stdtypes.html#dict)[[str](https://docs.python.org/3/builtins/stdtypes.html#str), [Any](https://docs.python.org/3/library/typing.html#typing.Any)][[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_learner_params.py#L163)

Return the hyper-parameters captured from the constructor.

#### predict(instance: [Instance](capymoa.core.Instance.md#capymoa.core.Instance)) → [int](https://docs.python.org/3/builtins/functions.html#int) | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_classifier.py#L56)

Predict the label of an instance.

The base implementation calls [`predict_proba()`](#capymoa.classifier.EFDT.predict_proba) and returns the
label with the highest probability.

* **Parameters:**
  **instance** – The instance to predict the label for.
* **Returns:**
  The predicted label or `None` if the classifier is unable
  to make a prediction.

#### predict_proba(instance) → [ndarray](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray)[[tuple](https://docs.python.org/3/builtins/stdtypes.html#tuple)[[Any](https://docs.python.org/3/library/typing.html#typing.Any), ...], [dtype](https://numpy.org/doc/stable/reference/generated/numpy.dtype.html#numpy.dtype)[float64]] | [None](https://docs.python.org/3/builtins/constants.html#None)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_classifier.py#L117)

Return probability estimates for each label.

* **Parameters:**
  **instance** – The instance to estimate the probabilities for.
* **Returns:**
  An array of probabilities for each label or `None` if the
  classifier is unable to make a prediction.

#### train(instance)[[source]](https://github.com/adaptive-machine-learning/CapyMOA/blob/3e255b1/src/capymoa/base/_classifier.py#L114)

Train the classifier with a labeled instance.

* **Parameters:**
  **instance** – The labeled instance to train the classifier with.

#### random_seed *: [int](https://docs.python.org/3/builtins/functions.html#int)*

The random seed for reproducibility.

When implementing a classifier ensure random number generators are seeded.

#### schema *: [Schema](capymoa.stream.Schema.md#capymoa.stream.Schema)*

The schema representing the instances.
