Falcun#
- class skactiveml.pool.Falcun(gamma=10, missing_label=nan, random_state=None, multilabel_aggregation_fn=<function mean>, target_type='auto')[source]#
Bases:
SingleAnnotatorPoolQueryStrategyFast Active Learning by Contrastive UNcertainty (FALCUN)
This class implements the “Fast Active Learning by Contrastive UNcertainty” (FALCUN) query strategy [1], which selects a batch directly in probability space using a self-adjusting mix of uncertainty and diversity. By operating only on low-dimensional class-probability outputs rather than deep embeddings, it achieves fast acquisitions while retaining strong label efficiency.
The distances in probability space are initialized with the uncertainty scores themselves (cf. Eq. (3) in [1]), so the first sample of a batch is sampled with a probability proportional to (2 * uncertainty) ** gamma and thus carries no diversity information. At batch_size=1, the acquisition is therefore gamma-tempered probabilistic margin sampling.
FALCUN was proposed for single-output classification. Multi-label support in this implementation is an extension and not part of the original proposal in [1]. For resolved multi-label targets, the paper’s top-two margin is applied to each label output independently, i.e., the per-label uncertainty of the label output j is the binary margin 1 - |2 * p_j - 1| of its positive-class probability p_j, and multilabel_aggregation_fn reduces these per-label margins along the label axis to the uncertainty of one sample. The diversity term stays the L1 distance in probability space, which for multi-label targets is taken between the independent per-output positive-class probabilities, i.e., sum_j |p_j(x) - p_j(x_query)|. Correlations between label outputs therefore influence neither term.
- Parameters:
- gammafloat > 0, default=10
Controls the randomness in the selection. A value of 0 corresponds to random sampling, while a value going to infinity corresponds to selecting the sample with the highest utility (relevance).
- missing_labelscalar or string or np.nan or None, default=np.nan
Value to represent a missing label.
- random_stateNone or int or np.random.RandomState, default=None
The random state to use.
- multilabel_aggregation_fncallable, default=np.mean
Callable reducing the per-label uncertainty scores of one sample to one uncertainty score. It is only used for resolved multi-label classification targets. It is called with the per-label scores of shape (n_samples, n_outputs) and the label axis passed as the axis keyword argument, and must return one score per sample within the range of that sample’s per-label scores, e.g. np.mean, np.average, np.median, np.min, np.max, or a quantile. np.sum is not supported, because its result grows with the number of label outputs. Only the callability of the reduction is validated at runtime, so a violating reduction silently changes the acquisition scale. Here, an inflated uncertainty would dominate the diversity term, which is min-max normalized to [0, 1] from the second selection of a batch onward.
- target_type“auto” or “single-output” or “multi-label”, default=”auto”
Declared target type. The strategy supports single-output and multi-label classification. A fitted classifier’s target specification is authoritative when available.
References
Methods
query(X, y, clf[, fit_clf, sample_weight, ...])Query the next samples to be labeled.
Get metadata routing of this object.
get_params([deep])Get parameters for this estimator.
set_params(**params)Set the parameters of this estimator.
- Falcun.query(X, y, clf, fit_clf=True, sample_weight=None, candidates=None, batch_size=1, return_utilities=False)[source]#
Query the next samples to be labeled.
- Parameters:
- Xarray-like of shape (n_samples, n_features)
Training data set, usually complete, i.e., including the labeled and unlabeled samples.
- yarray-like of shape (n_samples,) or (n_samples, n_outputs)
Labels of the training data set (possibly including unlabeled ones indicated by self.missing_label). For multi-label targets, a row y[i] must either contain only observed labels or only missing_label values, i.e., no mixing within a row. In this case, only multilabel classification problems, i.e. multiple binary classification tasks, are supported. predict_proba must then return either shape (n_samples, n_outputs) or a list of binary probability matrices with shape (n_samples, 2) per output.
- clfskactiveml.base.SkactivemlClassifier
Classifier implementing the methods fit and predict_proba.
- fit_clfbool, default=True
Defines whether the classifier clf should be fitted on X, y, and sample_weight.
- sample_weightarray-like of shape (n_samples,) or (n_samples, n_outputs), default=None
Weights of training samples in X. For two-dimensional y, one weight per sample is supported. Per-target weights are forwarded to clf.fit without additional validation and require estimator support.
- candidatesNone or array-like of shape (n_candidates), dtype=int or array-like of shape (n_candidates, n_features), default=None
If candidates is None, the unlabeled samples from (X,y) are considered as candidates.
If candidates is of shape (n_candidates,) and of type int, candidates is considered as the indices of the samples in (X,y).
If candidates is of shape (n_candidates, …), the candidate samples are directly given in candidates (not necessarily contained in X).
- batch_sizeint, default=1
The number of samples to be selected in one AL cycle.
- return_utilitiesbool, default=False
If true, also return the utilities based on the query strategy.
- Returns:
- query_indicesnumpy.ndarray of shape (batch_size)
The query indices indicate for which candidate sample a label is to be queried, e.g., query_indices[0] indicates the first selected sample.
If candidates is None or of shape (n_candidates,), the indexing refers to the samples in X.
If candidates is of shape (n_candidates, n_features), the indexing refers to the samples in candidates.
- utilitiesnumpy.ndarray of shape (batch_size, n_samples)
The utilities of samples after each selected sample of the batch, e.g., utilities[0] indicates the utilities used for selecting the first sample (with index query_indices[0]) of the batch. Utilities for labeled samples will be set to np.nan.
If candidates is None, the indexing refers to the samples in X.
If candidates is of shape (n_candidates,) and of type int, utilities refers to the samples in X.
If candidates is of shape (n_candidates, …), utilities refers to the indexing in candidates.
- Falcun.get_metadata_routing()#
Get metadata routing of this object.
Please check User Guide on how the routing mechanism works.
- Returns:
- routingMetadataRequest
A
MetadataRequestencapsulating routing information.
- Falcun.get_params(deep=True)#
Get parameters for this estimator.
- Parameters:
- deepbool, default=True
If True, will return the parameters for this estimator and contained subobjects that are estimators.
- Returns:
- paramsdict
Parameter names mapped to their values.
- Falcun.set_params(**params)#
Set the parameters of this estimator.
The method works on simple estimators as well as on nested objects (such as
Pipeline). The latter have parameters of the form<component>__<parameter>so that it’s possible to update each component of a nested object.- Parameters:
- **paramsdict
Estimator parameters.
- Returns:
- selfestimator instance
Estimator instance.
Examples using skactiveml.pool.Falcun#
Fast Active Learning by Contrastive UNcertainty (FALCUN)