Due: Tuesday, September 29 at 11:59 PM EDT
In this homework you will build a binary classifier using one categorical feature at a time, evaluate its predictions with a confusion matrix, and use ROC curves and area under the curve (AUC) to compare features.
By the end of the homework, you should be able to:
Credit: Adapted from materials by Thao Nguyen, based on materials created by Sara Mathieson and Allison Gong.
Download the Homework 4 starter files from Piazza. Your working folder should contain:
FeatureModel.py
Partition.py
run_roc.py
util.py
README.md
HW04_written_problems.pdf
HW04_written_problems.tex
data/
mushroom_train.arff
mushroom_test.arff
HW04 includes a written worksheet covering confusion matrices, evaluation metrics, threshold tradeoffs, ROC curves, AUC, model comparison, and runtime analysis. The starter materials include two versions:
HW04_written_problems.pdf, which you may print and complete by hand; andHW04_written_problems.tex, which you may edit to typeset your answers using
LaTeX.Choose one approach. Show your work and clearly label every final answer. You may use a calculator for arithmetic, but do not use Python or another programming language to solve the written problems.
If you complete the worksheet by hand, scan the finished pages into one clear,
legible PDF. If you use LaTeX, compile the completed .tex file to PDF. In
either case, name the finished file HW04_written_problems.pdf and include it
with your Gradescope submission.
Install the required packages if they are not already available:
python3 -m pip install numpy matplotlib
Run the complete analysis with:
python3 run_roc.py \
-r data/mushroom_train.arff \
-e data/mushroom_test.arff
Keep reusable calculations in functions or methods. Importing your files must
not run the full analysis. Call each file’s main() function only beneath an
if __name__ == "__main__": guard.
This homework uses the Mushroom Dataset from the UCI Machine Learning Repository. Each example is a mushroom described by 22 categorical features. The goal is to predict whether the mushroom is:
-1; or+1.You will create a separate model for each feature, compare the models using several evaluation metrics, and produce a final ROC plot containing the five most informative features.
The datasets use ARFF format. Like CSV, ARFF stores tabular data, but it also lists the feature names and their possible values in a header. For example:
@attribute 'cap-shape' {b, c, x, f, k, s}
For cap-shape, these abbreviations mean:
bell=b, conical=c, convex=x, flat=f, knobbed=k, sunken=s
The final attribute is the class label:
@attribute 'class' {e, p}
Here e means edible and p means poisonous.
The starter code provides:
Example, which stores one example’s feature values and label;Partition, which stores a dataset and its possible feature values;util.parse_args(), which reads the required command-line arguments; andutil.read_arff(), which reads an ARFF file into a Partition.In run_roc.py, call util.parse_args(), then use util.read_arff() to read
both datasets. Your program must accept:
-r/--train_filename: the training-data path; and-e/--test_filename: the test-data path.Verify that the files were read correctly. The expected dataset sizes are:
training examples: 6538
test examples: 1586
Read through Partition.py and util.py carefully. You may add helpers to
these files, but do not change the meaning of the existing classes or command-
line arguments.
Complete the FeatureModel class in FeatureModel.py.
The constructor receives a training Partition and the name of one feature.
For each possible value of that feature, calculate the fraction of matching
training examples whose label is positive (+1). Store these probabilities so
that classify() can use them later.
For example, the training data contains 355 bell-shaped mushrooms. Of these, 41 are poisonous, so the estimated probability that a bell-shaped mushroom is poisonous is:
41 / 355 = 0.1154929577
For cap-shape, your model should obtain probabilities close to:
{
"b": 0.11549295774647887,
"c": 1.0,
"x": 0.4652588555858311,
"f": 0.49667318982387476,
"k": 0.7259036144578314,
"s": 0.0,
}
Implement:
def classify(self, example, threshold):
...
Look up the positive-class probability associated with the example’s value for this model’s feature. Return:
+1 when the probability is greater than or equal to threshold; and-1 otherwise.Use the model to classify every example in the test set. Create a confusion matrix and calculate:
accuracy = (TP + TN) / (TP + TN + FP + FN)
false-positive rate = FP / (FP + TN)
true-positive rate = TP / (TP + FN)
Use helper functions or methods rather than repeating these calculations.
In FeatureModel.py, test a model using cap-shape and a threshold of 0.5.
The exact formatting is flexible, but your output should be readable and should
contain results equivalent to:
feature: cap-shape, threshold: 0.5
prediction
-1 1
----------------
actual -1 | 785 46
1 | 636 119
accuracy: 0.569987 (904 out of 1586 correct)
false-positive rate: 0.055355
true-positive rate: 0.157616
An ROC curve compares the true-positive rate (TPR), which should be high, with the false-positive rate (FPR), which should be low. Each point on the curve represents a different classification threshold.
Every ROC curve includes:
(0, 0), corresponding to classifying everything as negative; and(1, 1), corresponding to classifying everything as positive.Begin with one feature, such as cap-shape. Create a FeatureModel, evaluate
it across a range of thresholds, and record the resulting (FPR, TPR) pairs.
The following thresholds include values just outside [0, 1], ensuring that
the terminal points appear:
thresholds = np.linspace(-0.0001, 1.1, 20)
Plot false-positive rate on the x-axis and true-positive rate on the y-axis.
Using "o-" as the line style will show both the evaluated points and the lines
connecting them.
Use a loop to create one ROC curve for each feature and initially plot all of them on the same axes. Include a legend so you can identify each feature.
Visually inspect the curves and select the five features that come closest to
an ideal classifier. An ideal ROC curve rises toward (0, 1) while remaining
as far as possible above the diagonal from (0, 0) to (1, 1).
Hard-code the five selected feature names in a list, then generate the final plot using only those features. Save it as:
figures/roc_curve_top5.pdf
Your program must create figures/ when necessary. The final figure must have
a descriptive title, labeled axes, and a legend.
Implement your own method for approximating the area under each ROC curve. You may not use a built-in AUC function. Make sure the ROC points are ordered by false-positive rate before computing the area.
Use AUC to identify the five best features quantitatively. In README.md,
describe your algorithm, report the resulting features, and compare them with
the five you selected visually. You do not need to create a second figure for
the AUC-selected features.
Answer the following questions in README.md:
n training examples, p features per example, v
possible values for the selected feature, and a binary outcome. State your
assumptions and justify your answer.Complete the short, ungraded questionnaire at the end of README.md as well.
This homework uses one feature at a time. Describe how you could combine the
single-feature models into a more robust classification system. You may also
implement your idea. Document any additional code and results in README.md.
Submit one .zip file to the HW04 assignment on Gradescope. Preserve this
structure:
FeatureModel.py
Partition.py
run_roc.py
util.py
README.md
HW04_written_problems.pdf
figures/
roc_curve_top5.pdf
You do not need to include the data/ directory unless Gradescope’s
submission instructions explicitly request it.
Before submitting:
FeatureModel with cap-shape and threshold 0.5;figures/roc_curve_top5.pdf is regenerated;HW04_written_problems.pdf;README.md; andFollow the course collaboration and attribution policies. You may discuss the assignment at the level allowed by the course policy, but the submitted code and written analysis must be your own unless the assignment is explicitly designated as collaborative.