CSI 4106 - Fall 2026
Version: Sep 25, 2026 08:51
GridSearchCV to tune multiple pipelines using a consistent metric and explain when randomized search is useful.OpenML is an open platform for sharing datasets, algorithms, and experiments - to learn how to learn better, together.
Author: Vincent Sigillito
Source: Obtained from UCI
Please cite: UCI citation policy
Title: Pima Indians Diabetes Database
Sources:
Past Usage:
Smith,J.W., Everhart,J.E., Dickson,W.C., Knowler,W.C., & Johannes,R.S. (1988). Using the ADAP learning algorithm to forecast the onset of diabetes mellitus. In {it Proceedings of the Symposium on Computer Applications and Medical Care} (pp. 261–265). IEEE Computer Society Press.
The diagnostic, binary-valued variable investigated is whether the patient shows signs of diabetes according to World Health Organization criteria (i.e., if the 2 hour post-load plasma glucose was at least 200 mg/dl at any survey examination or if found during routine medical care). The population lives near Phoenix, Arizona, USA.
Results: Their ADAP algorithm makes a real-valued prediction between 0 and 1. This was transformed into a binary decision using a cutoff of 0.448. Using 576 training instances, the sensitivity and specificity of their algorithm was 76% on the remaining 192 instances.
Relevant Information: Several constraints were placed on the selection of these instances from a larger database. In particular, all patients here are females at least 21 years old of Pima Indian heritage. ADAP is an adaptive learning routine that generates and executes digital analogs of perceptron-like devices. It is a unique algorithm; see the paper for details.
Number of Instances: 768
Number of Attributes: 8 plus class
For Each Attribute: (all numeric-valued)
Missing Attribute Values: None
Class Distribution: (class value 1 is interpreted as “tested positive for diabetes”)
Class Value Number of instances 0 500 1 268
Brief statistical analysis:
Attribute number: Mean: Standard Deviation:
3.8 3.4 120.9 32.0 69.1 19.4 20.5 16.0 79.8 115.2 32.0 7.9 0.5 0.3 33.2 11.8Relabeled values in attribute ‘class’ From: 0 To: tested_negative
From: 1 To: tested_positive
Downloaded from openml.org.
return_X_yfetch_openml returns a Bunch by default, or X and y when return_X_y=True.
The positive class is less frequent, but both classes have enough examples for stratified cross-validation.
Following the convention used by the corrected PimaIndiansDiabetes2 version of this dataset, zero values in five clinical measurements are treated as missing. Zero remains valid for the number of pregnancies.
This division is often called the holdout method.
Common starting point: Allocate 80% of the data to training and reserve 20% for testing.
Training set: Used to fit models and make model-building choices.
Test set: An independent subset used once for the final evaluation.
Training error: The prediction error measured on examples used to fit the model.
Generalization error: The expected prediction error on new examples drawn under similar conditions.
It cannot be observed directly. Validation and test scores provide estimates from finite samples.
Underfitting:
Overfitting:
Cross-validation estimates how well a model-building procedure will perform on unseen examples.
It repeatedly partitions the training data, trains a new classifier on some folds, and evaluates that classifier on the remaining fold. The test set remains untouched.
import matplotlib.pyplot as plt
def plot_k_fold_cross_validation(k):
fig, ax = plt.subplots()
matrix = np.ones((k, k))
np.fill_diagonal(matrix, 0)
ax.imshow(matrix, cmap='binary', interpolation='none')
for i in range(k):
for j in range(k):
label = 'Validation' if i == j else 'Train'
color = "black" if i == j else "white"
ax.text(j, i, label, ha="center", va="center", color=color)
ax.set_xticks(np.arange(k))
ax.set_yticks(np.arange(k))
ax.set_xticklabels([f"Fold {i+1}" for i in range(k)])
ax.set_yticklabels([f"Iteration {i+1}" for i in range(k)])
plt.title(f'{k}-Fold Cross-validation')
ax.grid(False)
plt.show()Data leakage occurs when information from a validation fold influences the procedure fitted on the remaining folds.
The validation examples influence the means and scales used to evaluate them.
fit_transform and transformfit(X) learns information from X.fit_transform(X) learns that information and applies the transformation to the same examples.transform(X) applies the stored information without recomputing it.Validation, test, and future examples receive transform, never fit or fit_transform.
For every fold, cross-validation creates and fits a new pipeline.
| Pipeline step | Fitting folds | Validation fold |
|---|---|---|
| Imputer | fit_transform |
transform |
| Scaler | fit_transform |
transform |
| Classifier | fit |
predict_proba / decision_function |
Learned medians, means, and scales are not carried into the next fold.
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier
tree_pipeline = Pipeline([
("imputer", SimpleImputer(strategy="median")),
("model", DecisionTreeClassifier(random_state=seed)),
])
knn_pipeline = Pipeline([
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler()),
("model", KNeighborsClassifier()),
])
logistic_pipeline = Pipeline([
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler()),
("model", LogisticRegression(
solver="saga",
max_iter=3000,
random_state=seed,
)),
])cross_val_scoreScores: [0.81133721 0.76061047 0.84447674 0.74883721 0.75639881]
Mean AUROC: 0.784
Standard deviation: 0.037
Model parameters are learned from the training data.
Hyperparameters are chosen before each fit and control the model structure or learning process.
criterion: chooses how the quality of a split is measured: gini, entropy, or log_loss.max_depth: limits the maximum depth of the tree.C: inverse regularization strength. Smaller values apply stronger regularization.l1_ratio: selects the balance between L2 and L1 regularization when the saga solver is used.n_neighbors: number of neighbours used to make a prediction.weights: gives neighbours equal influence (uniform) or greater influence when they are closer (distance).n_neighborsneighbor_scores = {}
neighbor_values = [1, 3, 5, 7, 9, 11, 15, 21, 31]
for value in neighbor_values:
candidate = knn_pipeline.set_params(
model__n_neighbors=value,
model__weights="uniform",
)
scores = cross_val_score(
candidate,
X_train,
y_train,
cv=cv,
scoring=selection_metric,
)
neighbor_scores[value] = scores.mean()
print(f"n_neighbors={value:2d}: {scores.mean():.3f}")
best_n_neighbors = max(neighbor_scores, key=neighbor_scores.get)
print("Best n_neighbors:", best_n_neighbors)n_neighborsn_neighbors= 1: 0.664
n_neighbors= 3: 0.764
n_neighbors= 5: 0.784
n_neighbors= 7: 0.798
n_neighbors= 9: 0.815
n_neighbors=11: 0.822
n_neighbors=15: 0.832
n_neighbors=21: 0.831
n_neighbors=31: 0.831
Best n_neighbors: 15
weightsweight_scores = {}
for value in ["uniform", "distance"]:
candidate = knn_pipeline.set_params(
model__n_neighbors=best_n_neighbors,
model__weights=value,
)
scores = cross_val_score(
candidate,
X_train,
y_train,
cv=cv,
scoring=selection_metric,
)
weight_scores[value] = scores.mean()
print(f"weights={value:8s}: {scores.mean():.3f}")weights=uniform : 0.832
weights=distance: 0.833
Several hyperparameters may need to be considered together.
Manual exploration of their combinations is tedious and error-prone.
Grid search evaluates a predefined set of combinations systematically:
Enumerate the Cartesian product of the candidate values.
Evaluate each combination using the same cross-validation folds and metric.
GridSearchCV: decision treefrom sklearn.model_selection import GridSearchCV
tree_param_grid = {
"model__max_depth": [2, 3, 4, 5, 6, None],
"model__criterion": ["gini", "entropy", "log_loss"],
}
tree_search = GridSearchCV(
tree_pipeline,
tree_param_grid,
cv=cv,
scoring=selection_metric,
)
tree_search.fit(X_train, y_train)
(tree_search.best_params_, tree_search.best_score_)({'model__criterion': 'gini', 'model__max_depth': 4},
np.float64(0.7881160022148395))
GridSearchCV: KNN({'model__n_neighbors': 21, 'model__weights': 'distance'},
np.float64(0.8336620985603543))
GridSearchCV: logistic regression({'model__C': 0.1, 'model__l1_ratio': 0.0}, np.float64(0.8443867663344408))
model_searches = {
"Decision tree": tree_search,
"KNN": knn_search,
"Logistic regression": logistic_search,
}
for name, search in model_searches.items():
print(f"{name:20s}: {search.best_score_:.3f}")
best_model_name = max(
model_searches,
key=lambda name: model_searches[name].best_score_,
)
best_search = model_searches[best_model_name]
print("Selected model:", best_model_name)Decision tree : 0.788
KNN : 0.834
Logistic regression : 0.844
Selected model: Logistic regression
RandomizedSearchCV samples candidates from supplied values or probability distributions.from sklearn.metrics import classification_report, roc_auc_score
best_model = best_search.best_estimator_
y_test_score = best_model.predict_proba(X_test)[:, 1]
y_test_pred = best_model.predict(X_test)
print("Selected model:", best_model_name)
print(f"Test AUROC: {roc_auc_score(y_test, y_test_score):.3f}")
print(classification_report(y_test, y_test_pred))Selected model: Logistic regression
Test AUROC: 0.810
precision recall f1-score support
0 0.74 0.80 0.77 100
1 0.57 0.48 0.52 54
accuracy 0.69 154
macro avg 0.65 0.64 0.64 154
weighted avg 0.68 0.69 0.68 154
GridSearchCV to tune and compare decision tree, KNN, and logistic regression pipelines with the same folds and AUROC metric. Randomized search provides an alternative for larger search spaces.Marcel Turcotte
School of Electrical Engineering and Computer Science (EECS)
University of Ottawa