UHerding#
- class skactiveml.pool.UHerding(method='margin_sampling', predict_proba_dict=None, predict_proba_parser=None, temperatures=None, validation_size=0.2, n_ece_bins=15, normalize_samples=True, metric='rbf', metric_dict=None, adaptive_sigma=True, missing_label=nan, random_state=None, multilabel_aggregation_fn=<function mean>, target_type='auto')[source]#
Bases:
SingleAnnotatorPoolQueryStrategyUncertainty Herding (UHerding)
“Uncertainty Herding” (UHerding) is a query strategy [1] that greedily maximizes an uncertainty-weighted coverage objective in feature space. In addition to the greedy selection itself, the implementation follows the parameter adaptation scheme of the paper:
select a temperature based on calibration via train/validation splits of the currently labeled set,
adapt the Gaussian kernel radius to the current labeled feature space.
UHerding was proposed for single-output classification. Multi-label support in this implementation is an extension and not part of the original proposal in [1]. For resolved multi-label targets, the temperature-scaled probabilities are the per-output sigmoids expit(logits / tau) with the per-output temperature tau instead of a softmax, whenever logits are available, and the per-label uncertainty of the label output j is computed by method from its positive-class probability p_j alone (cf. uncertainty_scores). multilabel_aggregation_fn reduces these per-label scores along the label axis to the uncertainty weight of one sample, which then scales that sample’s coverage gains. The coverage objective itself is unchanged, because it operates on the sample representations only. Correlations between label outputs therefore influence neither the uncertainty weight nor the coverage objective.
- Parameters:
- method‘least_confident’ or ‘margin_sampling’ or ‘entropy’, default=’margin_sampling’
Uncertainty definition applied to temperature-scaled probabilities.
- predict_proba_dictdict or None, default=None
Optional keyword arguments forwarded to clf.predict_proba to request additional outputs such as logits and embeddings.
If predict_proba_parser is None, optional outputs are interpreted by the default convention (probas, logits, embeddings). Typical usage with SkorchClassifier is therefore:
predict_proba_dict={"extra_outputs": ["logits", "emb"]}
If logits are not returned by predict_proba, decision_function is used as a fallback when available, e.g. for scikit-learn logistic regression models wrapped by SklearnClassifier.
- predict_proba_parsercallable or None, default=None
Optional parser applied to the raw return value of clf.predict_proba(X, **predict_proba_dict).
The parser must return either (probas, logits) or (probas, logits, embeddings). probas may be None, in which case they are computed from logits via softmax. embeddings may be None, in which case the original samples are used.
If None, the default convention is used:
array output: treated as probas,
tuple output: treated as (probas, logits, embeddings).
- temperaturesfloat or array-like of shape (n_temperatures,) or None, default=None
Temperature specification for calibrating the uncertainty estimates. Two behaviors are distinguished:
A positive float or a length-one array-like is a fixed temperature shared by every output. It is used directly, i.e. no calibration train/validation split and no internal calibration refit take place.
An array-like with more than one entry is a shared candidate grid searched by expected calibration error on a validation split of the currently labeled set. For a single-output target, one temperature is selected. For a multi-label target, the grid is searched independently per label output, yielding one selected temperature per output.
If None, the candidate grid np.logspace(-1, 1, 49) is used.
- validation_sizefloat or int, default=0.2
Validation size passed to the calibration train/validation split.
- n_ece_binsint, default=15
Number of bins used for the expected calibration error.
- normalize_samplesbool, default=True
Flag whether to normalize feature vectors to unit length before computing pairwise distances and kernels.
- metricstr or callable, default=’rbf’
Kernel used for the coverage objective.
- metric_dictdict or None, default=None
Optional keyword arguments passed to pairwise_kernels.
- adaptive_sigmabool, default=True
Flag whether to adapt the radius according to the minimum non-zero labeled pairwise distance. This option requires metric=’rbf’.
- missing_labelscalar or string or np.nan or None, default=np.nan
Value to represent a missing label.
- random_stateNone or int or np.random.RandomState, default=None
The random state to use.
- multilabel_aggregation_fncallable, default=np.mean
Callable reducing the per-label uncertainty scores of one sample to one uncertainty weight. It is only used for resolved multi-label targets. It is called with the per-label scores of shape (n_samples, n_outputs) and the label axis passed as the axis keyword argument, and must return one score per sample within the range of that sample’s per-label scores, e.g. np.mean, np.average, np.median, np.min, np.max, or a quantile. np.sum is not supported, because its result grows with the number of label outputs. Only the callability of the reduction is validated at runtime, so a violating reduction silently changes the acquisition scale.
- target_type“auto” or “single-output” or “multi-label”, default=”auto”
Declared target type. The strategy supports single-output and multi-label classification. A fitted classifier’s target specification is authoritative when available.
References
Methods
query(X, y, clf[, fit_clf, sample_weight, ...])Determines for which candidate samples labels are to be queried.
Get metadata routing of this object.
get_params([deep])Get parameters for this estimator.
set_params(**params)Set the parameters of this estimator.
- UHerding.query(X, y, clf, fit_clf=True, sample_weight=None, candidates=None, batch_size=1, return_utilities=False)[source]#
Determines for which candidate samples labels are to be queried.
- Parameters:
- Xarray-like of shape (n_samples, …)
Training data set, usually complete, i.e., including the labeled and unlabeled samples.
- yarray-like of shape (n_samples,) or (n_samples, n_outputs)
Labels of the training data set (possibly including unlabeled ones indicated by self.missing_label). For multi-label targets, a row y[i] must either contain only observed labels or only missing_label values, i.e., no mixing within a row.
- clfskactiveml.base.SkactivemlClassifier
Classifier implementing fit and predict_proba. For temperature-scaled uncertainty estimation, the classifier should either provide logits via predict_proba extras or implement decision_function. Otherwise, the non-calibrated probabilities are used as fallback. For multilabel classification, predict_proba must return either shape (n_samples, n_outputs) or a list of binary probability matrices with shape (n_samples, 2) per output.
- fit_clfbool, default=True
Defines whether the classifier clf should be fitted on X, y, and sample_weight before evaluating the acquisition function. Independent of this flag, temporary cloned classifiers may still be fitted internally to select the temperature parameter.
- sample_weightarray-like of shape (n_samples,) or (n_samples, n_outputs), default=None
Weights of training samples in X. For two-dimensional y, one weight per sample is supported. Per-target weights are forwarded to clf.fit without additional validation and require estimator support.
- candidatesNone or array-like of shape (n_candidates,), dtype=int or array-like of shape (n_candidates, …), default=None
If candidates is None, the unlabeled samples from (X, y) are considered as candidates.
If candidates is of shape (n_candidates,) and of type int, candidates is considered as the indices of the samples in (X, y).
If candidates is of shape (n_candidates, …), the candidate samples are directly given in candidates (not necessarily contained in X).
- batch_sizeint, default=1
The number of samples to be selected in one AL cycle.
- return_utilitiesbool, default=False
If True, also return the utilities based on the query strategy.
- Returns:
- query_indicesnumpy.ndarray of shape (batch_size,)
The query indices indicate for which candidate sample a label is to be queried, e.g., query_indices[0] indicates the first selected sample.
If candidates is None or of shape (n_candidates,), the indexing refers to the samples in X.
If candidates is of shape (n_candidates, …), the indexing refers to the samples in candidates.
- utilitiesnumpy.ndarray of shape (batch_size, n_samples) or numpy.ndarray of shape (batch_size, n_candidates)
The utilities of samples after each selected sample of the batch, e.g., utilities[0] indicates the utilities used for selecting the first sample (with index query_indices[0]) of the batch. Utilities for labeled samples or already selected candidates are set to np.nan.
If candidates is None, the indexing refers to the samples in X.
If candidates is of shape (n_candidates,) and of type int, utilities refers to the samples in X.
If candidates is of shape (n_candidates, …), utilities refers to the indexing in candidates.
- UHerding.get_metadata_routing()#
Get metadata routing of this object.
Please check User Guide on how the routing mechanism works.
- Returns:
- routingMetadataRequest
A
MetadataRequestencapsulating routing information.
- UHerding.get_params(deep=True)#
Get parameters for this estimator.
- Parameters:
- deepbool, default=True
If True, will return the parameters for this estimator and contained subobjects that are estimators.
- Returns:
- paramsdict
Parameter names mapped to their values.
- UHerding.set_params(**params)#
Set the parameters of this estimator.
The method works on simple estimators as well as on nested objects (such as
Pipeline). The latter have parameters of the form<component>__<parameter>so that it’s possible to update each component of a nested object.- Parameters:
- **paramsdict
Estimator parameters.
- Returns:
- selfestimator instance
Estimator instance.