max_loss_reduction_max_confidence#

skactiveml.pool.max_loss_reduction_max_confidence(probas, n_positive_labels)[source]#

Calculate the maximum loss reduction with maximal confidence.

For each candidate sample, the n_positive_labels most probable labels are predicted positive and the remaining ones negative [1]. The loss reduction of this most confident labeling is the sum of the hinge-style losses (1 - yhat * (2 * probas - 1)) / 2, i.e., it sums 1 - probas for the labels predicted positive and probas for the labels predicted negative.

Parameters:
probasarray-like of shape (n_candidates, n_outputs)

Canonical positive-class probabilities of the candidate samples, i.e., one probability per output. The equivalent list of (n_candidates, 2) binary probability matrices is canonicalized as well, although query strategies are expected to canonicalize at their own boundary.

n_positive_labelsarray-like of shape (n_candidates,)

Predicted number of positive labels per candidate sample, e.g., as predicted by a label-cardinality discriminator. Each entry must be an integer in [0, n_outputs].

Returns:
utilitiesnumpy.ndarray of shape (n_candidates,)

Loss reduction of each candidate sample under its most confident labeling, i.e., one finite value in [0, n_outputs] per candidate. Larger values indicate more useful candidates.

Raises:
ValueError

If n_positive_labels is not a one-dimensional array of integers within [0, n_outputs], if probas is not a multi-label probability matrix with one row per entry of n_positive_labels, or if probas contains values outside of [0, 1].

References

[1]

Yang, B., Sun, J.-T., Wang, T., & Chen, Z. (2009). Effective Multi-Label Active Learning for Text Classification. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 917-926).