uncertainty_scores#

skactiveml.pool.uncertainty_scores(probas, cost_matrix=None, method='least_confident', is_multilabel=False, multilabel_aggregation_fn=<function mean>)[source]#

Computes uncertainty scores. Three methods are available: least confident (‘least_confident’), margin sampling (‘margin_sampling’), and entropy based uncertainty (‘entropy’) [1]. For the least confident and margin sampling methods cost-sensitive variants are implemented in case of a given cost matrix (see [2] for more information). For multilabel data, only ‘least_confident’, ‘margin_sampling’, and ‘entropy’ are supported.

The three uncertainty measures were proposed for single-output classification. Multi-label support in this implementation is an extension and not part of the original proposal in [1]. It decomposes the target per label output, i.e., the per-label score of the label output j is computed from its positive-class probability p_j alone,

  • min(p_j, 1 - p_j) for ‘least_confident’,

  • 1 - |2 * p_j - 1| for ‘margin_sampling’, i.e., the margin between the two classes of the label output, and

  • -(p_j * log(p_j) + (1 - p_j) * log(1 - p_j)) for ‘entropy’, i.e., the binary entropy of the label output, with the endpoint terms 0 * log(0) defined as zero,

and multilabel_aggregation_fn then reduces these per-label scores along the label axis to one score per sample. Correlations between label outputs are ignored by construction.

Parameters:
probasarray-like of shape (n_samples, n_classes)

Class membership probabilities for each sample. If is_multilabel=True, positive-class probabilities of shape (n_samples, n_outputs) or a list of binary probability matrices with shape (n_samples, 2) per label output are expected instead.

cost_matrixarray-like pf shape (n_classes, n_classes)

Cost matrix with cost_matrix[i,j] defining the cost of predicting class j for a sample with the actual class i. Only supported for ‘least_confident’ or ‘margin_sampling’ and single-output targets. Cost matrices are not supported if is_multilabel=True.

method‘least_confident’ or ‘margin_sampling’ or ‘entropy’, default=’least_confident’

The method to calculate the uncertainty. For multilabel data, only ‘least_confident’, ‘margin_sampling’, and ‘entropy’ are supported.

is_multilabelbool, default=False

Flag whether probas are multi-label positive-class probabilities.

multilabel_aggregation_fncallable, default=np.mean

Callable reducing the per-label uncertainty scores of one sample to one uncertainty score. It is only used if is_multilabel=True. It is called with the per-label scores of shape (n_samples, n_outputs) and the label axis passed as the axis keyword argument, and must return one score per sample within the range of that sample’s per-label scores, e.g. np.mean, np.average, np.median, np.min, np.max, or a quantile. np.sum is not supported, because its result grows with the number of label outputs. Only the callability of the reduction is validated at runtime, so a violating reduction silently changes the acquisition scale.

References

[1] (1,2)

Settles, Burr. “Active learning literature survey”. University of Wisconsin-Madison Department of Computer Sciences, 2009.

[2]

P.-L. Chen and H.-T. Lin. Active Learning for Multiclass Cost-Sensitive Classification Using Probabilistic Models. In Conf. Technol. Appl. Artif. Intell., pages 13–18, 2013.