SklearnClassifier#
- class skactiveml.classifier.SklearnClassifier(estimator, include_unlabeled_samples=False, classes=None, missing_label=nan, cost_matrix=None, random_state=None, proba_format='auto', target_type='auto')[source]#
Bases:
SkactivemlClassifier,MetaEstimatorMixinSklearn Classifier
Implementation of a wrapper class for scikit-learn [1] classifiers such that
missing labels can be handled, e.g., by filtering them,
classes can be fixed at initialization, e.g., to have consistent probabilistic outputs even when there are no observed labels for each class,
cost-sensitive decisions can be made, e.g., to consider different types of misclassification costs.
- Parameters:
- estimatorsklearn.base.ClassifierMixin
The scikit-learn classifier to be wrapped. A predict_proba method is only required when predict_proba or cost_matrix based prediction is used.
- include_unlabeled_samplesbool, default=False
If False, only labeled samples are passed to the fit method of the estimator.
If True, all samples including the unlabeled ones are passed to the fit method of the estimator. Ensure that your estimator is able to handle unlabeled samples marked by missing_label. Otherwise, missing_label is interpreted as a regular class label. Note that semi-supervised classifiers of sklearn expect missing_label=-1.
- classesarray-like of shape (n_classes,), or a list of such array-likes, default=None
A flat vocabulary describes single-output classification.
Nested binary vocabularies describe multi-label classification, one class vocabulary per label output. With explicit target_type=”multi-label”, vocabularies can instead be resolved from y when classes=None and both classes are observed in every output.
- missing_labelscalar or string or np.nan or None, default=np.nan
Value to represent a missing label.
- cost_matrixarray-like of shape (n_classes, n_classes)
Cost matrix with cost_matrix[i,j] indicating cost of predicting class classes[j] for a sample of class classes[i]. Can be only set, if classes is not None and in the case of single output problems.
- random_stateint or RandomState instance or None, default=None
Determines random number for predict method. Pass an int for reproducible results across multiple method calls.
- proba_format“auto” or “list” or “array”, default=”auto”
- Output format of ``predict_proba``.
- - Single-output: always returns an array of shape `(n_samples, n_classes)`.
- - Multilabel (2D targets with binary classes per output):
‘list’ -> list of (n_samples, 2) arrays
‘array’ -> array of shape (n_samples, n_outputs) with P(y=pos_label)
- target_type“auto” or “single-output” or “multi-label”, default=”auto”
Declared target type. Single-output classification is always supported. Multi-label classification requires an estimator that implements predict_proba and positively declares either target_tags.multi_output or classifier_tags.multi_label, e.g., sklearn.multioutput.MultiOutputClassifier, sklearn.multiclass.OneVsRestClassifier, or sklearn.ensemble.RandomForestClassifier. An explicit “multi-label” can resolve binary per-label vocabularies from observed targets when classes is None.
- Attributes:
- target_spec_skactiveml.utils.TargetSpec
Immutable target specification established by a successful fit. Use its classes field for canonical class ordering.
Notes
A pre-fitted estimator already published the target semantics of its own predictions, so its learned classes are reconciled with the declared ones by class identity before any fitted attribute of this wrapper is published. Declared classes may extend the learned class vocabulary, whose additional classes then receive zero-filled probability columns in the declared order, but they can neither reinterpret learned classes nor change the number of predicted outputs. A flat classes_ published together with fitted multi-label metadata, e.g., by sklearn.multiclass.OneVsRestClassifier, identifies binary indicator outputs instead of describing one class vocabulary, so one binary indicator vocabulary per output has to be declared through classes. A pre-fitted estimator publishing neither one class vocabulary per label output nor such metadata cannot be declared multi-label, because its flat learned vocabulary is indistinguishable from single-output classification.
Attributes this wrapper does not hold itself are read from the wrapped estimator. The fitted attributes it resolves itself, e.g. classes_ and target_spec_, never are: around a pre-fitted estimator they resolve this wrapper’s own target semantics on first access, and they raise the usual not-fitted error while no such semantics exist. A pre-fitted estimator’s learned classes stay readable as estimator.classes_, which states whose vocabulary they are.
Two degenerate training cases are part of this wrapper’s contract and make it predict the observed class label distribution instead of raising: an empty labeled training subset, and an estimator rejecting a labeled training subset that carries fewer than two distinct classes in at least one output. Both set is_fitted_ to False and emit a warning. Every other estimator failure is raised.
References
[1]Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 12, 2011, 2825–2830.
Methods
fit(X[, y, sample_weight])Fit the model using X as training data and y as class labels.
partial_fit(X[, y, sample_weight])Partially fitting the model using X as training data and y as class labels.
predict(X, **predict_kwargs)Return class label predictions for the input data X.
predict_proba(X, **predict_proba_kwargs)Return probability estimates for the input data X.
score(X, y[, sample_weight])Return the mean accuracy on the given test data and labels.
Get metadata routing of this object.
get_params([deep])Get parameters for this estimator.
set_fit_request(*[, sample_weight])Configure whether metadata should be requested to be passed to the
fitmethod.set_params(**params)Set the parameters of this estimator.
set_partial_fit_request(*[, sample_weight])Configure whether metadata should be requested to be passed to the
partial_fitmethod.set_score_request(*[, sample_weight])Configure whether metadata should be requested to be passed to the
scoremethod.
- SklearnClassifier.fit(X, y=None, sample_weight=None, **fit_kwargs)[source]#
Fit the model using X as training data and y as class labels.
- Parameters:
- Xarray-like of shape (n_samples, …)
The feature matrix representing the samples.
- yarray-like of shape (n_samples,) or (n_samples, n_outputs)
Labels of the training data set (possibly including unlabeled ones indicated by missing_label). For multilabel problems, a row y[i] must either contain only observed labels or only missing_label values, i.e., no mixing within a row. Note that Y (capitalized) is only accepted if the wrapped estimator exposes this parameter name in its fit signature.
- sample_weightarray-like of shape (n_samples,) or (n_samples, n_outputs)
It contains the weights of the training samples’ class labels. Only supported if the wrapped sklearn classifier can handle sample weights.
- fit_kwargsdict-like
Further parameters as input to the fit method of the estimator.
- Returns:
- self: SklearnClassifier,
The SklearnClassifier object fitted on the training data.
- SklearnClassifier.partial_fit(X, y=None, sample_weight=None, **fit_kwargs)[source]#
Partially fitting the model using X as training data and y as class labels.
- Parameters:
- Xarray-like of shape (n_samples, …)
The feature matrix representing the samples.
- yarray-like of shape (n_samples,) or (n_samples, n_outputs)
Labels of the training data set (possibly including unlabeled ones indicated by missing_label). For multilabel problems, a row y[i] must either contain only observed labels or only missing_label values, i.e., no mixing within a row. Note that Y (capitalized) is only accepted if the wrapped estimator exposes this parameter name in its partial_fit signature.
- sample_weightarray-like of shape (n_samples,) or (n_samples, n_outputs)
It contains the weights of the training samples’ class labels. Only supported if the wrapped sklearn classifier can handle sample weights.
- fit_kwargsdict-like
Further parameters as input to the partial_fit method of the estimator.
- Returns:
- selfSklearnClassifier,
The SklearnClassifier object fitted on the training data.
- SklearnClassifier.predict(X, **predict_kwargs)[source]#
Return class label predictions for the input data X.
- Parameters:
- Xarray-like of shape (n_samples, …)
Input samples.
- predict_kwargsdict-like
Further parameters as input to the predict method of the estimator.
- Returns:
- y_prednumpy.ndarray of shape (n_samples,) or (n_samples, n_outputs)
Predicted class labels of the input samples.
- SklearnClassifier.predict_proba(X, **predict_proba_kwargs)[source]#
Return probability estimates for the input data X.
- Parameters:
- Xarray-like of shape (n_samples, …)
Input samples.
- predict_proba_kwargsdict-like
Further parameters as input to the predict_proba method of the estimator.
- Returns:
- Pnumpy.ndarray of shape (n_samples, n_classes), numpy.ndarray of shape (n_samples, n_outputs), or list of numpy.ndarray
The class probabilities of the input samples. For single-output classification, the return value has shape (n_samples, n_classes). For multilabel classification, proba_format=’array’ returns shape (n_samples, n_outputs) with positive-class probabilities and proba_format=’list’ returns one (n_samples, 2) array per output.
- SklearnClassifier.score(X, y, sample_weight=None)#
Return the mean accuracy on the given test data and labels.
- Parameters:
- Xarray-like of shape (n_samples, …)
Test samples.
- yarray-like of shape (n_samples,)
True class labels of the test samples X.
- sample_weightarray-like of shape (n_samples,), default=None
Sample weights of the test sample X.
- Returns:
- scorefloat
Mean accuracy of self.predict(X) regarding y.
- SklearnClassifier.get_metadata_routing()#
Get metadata routing of this object.
Please check User Guide on how the routing mechanism works.
- Returns:
- routingMetadataRequest
A
MetadataRequestencapsulating routing information.
- SklearnClassifier.get_params(deep=True)#
Get parameters for this estimator.
- Parameters:
- deepbool, default=True
If True, will return the parameters for this estimator and contained subobjects that are estimators.
- Returns:
- paramsdict
Parameter names mapped to their values.
- SklearnClassifier.set_fit_request(*, sample_weight: bool | None | str = '$UNCHANGED$') SklearnClassifier#
Configure whether metadata should be requested to be passed to the
fitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed tofitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it tofit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
- sample_weightstr, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED
Metadata routing for
sample_weightparameter infit.
- Returns:
- selfobject
The updated object.
- SklearnClassifier.set_params(**params)#
Set the parameters of this estimator.
The method works on simple estimators as well as on nested objects (such as
Pipeline). The latter have parameters of the form<component>__<parameter>so that it’s possible to update each component of a nested object.- Parameters:
- **paramsdict
Estimator parameters.
- Returns:
- selfestimator instance
Estimator instance.
- SklearnClassifier.set_partial_fit_request(*, sample_weight: bool | None | str = '$UNCHANGED$') SklearnClassifier#
Configure whether metadata should be requested to be passed to the
partial_fitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed topartial_fitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it topartial_fit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
- sample_weightstr, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED
Metadata routing for
sample_weightparameter inpartial_fit.
- Returns:
- selfobject
The updated object.
- SklearnClassifier.set_score_request(*, sample_weight: bool | None | str = '$UNCHANGED$') SklearnClassifier#
Configure whether metadata should be requested to be passed to the
scoremethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed toscoreif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it toscore.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
- sample_weightstr, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED
Metadata routing for
sample_weightparameter inscore.
- Returns:
- selfobject
The updated object.
Examples using skactiveml.classifier.SklearnClassifier#
Batch Bayesian Active Learning by Disagreement (BatchBALD)