Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Scores, Thresholds, and ROC Curves

The previous section evaluated classifiers after the decision was already made: the system assigned a label, and we counted outcomes. But most classifiers do not jump directly to a label. They produce a continuous score, a confidence value between 0 and 1, and the label emerges only after comparing that score to a threshold. The spam filter used later in this section assigns 0.90 to obvious junk, 0.10 to the legitimate email “Lunch plans?”, and 0.40 to the ambiguous spam message “Claim your reward points”. At threshold 0.5, the latter message stays in the inbox; at threshold 0.4, it moves to junk.

Different thresholds produce different confusion matrices, different precision/recall values, and different TPR/FPR trade-offs. A single classifier is not one point in metric space; it is a family of points, one for each threshold. This section explores how to visualize that family and choose a single operating point.

Score Distributions and the Threshold

A binary classifier assigns a score s(x)∈[0,1]s(x) \in [0, 1] to each item xx. Items from the positive class tend to receive higher scores; items from the negative class tend to receive lower scores. If the two score distributions were perfectly separated (all positives above some value, all negatives below), any threshold between them would classify perfectly. In practice, the distributions overlap. Figure 1 visualizes how that overlap creates unavoidable false positives and false negatives.

Score distributions for the negative class (grey) and positive class (yellow) with a decision threshold. The overlap region produces classification errors regardless of where the threshold is placed.

Figure 1:Score distributions for the negative class (grey) and positive class (yellow) with a decision threshold. The overlap region produces classification errors regardless of where the threshold is placed.

The threshold TT partitions the score axis into two regions: items with s(x)≥Ts(x) \geq T are predicted positive, items with s(x)<Ts(x) < T are predicted negative. Moving the threshold changes the balance between true positives, false positives, true negatives, and false negatives:

This is the fundamental trade-off underlying every threshold-based classifier. No single threshold is universally best; the right choice depends on the error costs of the application.

Formalizing the Threshold as Areas Under the Distributions

Let fn(x)f_n(x) denote the score distribution of the negative class and fp(x)f_p(x) the score distribution of the positive class. Each classification rate corresponds to an area under one of these curves, partitioned by the threshold TT. Figure 2 shows the four resulting areas:

Decomposition of the score distributions into the four classification rates. TNR and FPR partition the negative-class curve f_n(x); FNR and TPR partition the positive-class curve f_p(x).

Figure 2:Decomposition of the score distributions into the four classification rates. TNR and FPR partition the negative-class curve fn(x)f_n(x); FNR and TPR partition the positive-class curve fp(x)f_p(x).

The integrals make the trade-off precise. Shifting TT to the left expands the integration range for TPR (the area under fpf_p to the right of TT grows), but simultaneously expands the integration range for FPR (the area under fnf_n to the right of TT also grows). No threshold eliminates the overlap region; it can only choose how to distribute the overlap errors between false positives and false negatives.

The ROC Curve

Instead of evaluating one threshold, the ROC (Receiver Operating Characteristic) curve evaluates all of them simultaneously. It plots TPR(T)TPR(T) on the vertical axis against FPR(T)FPR(T) on the horizontal axis, sweeping TT from high to low. The result is a curve from (0,0)(0, 0) to (1,1)(1, 1) that shows the full trade-off landscape.

Figure 3 shows 20 emails (10 spam, 10 legitimate) sorted by descending spam score. At each score, we set the threshold there and count cumulative TP, FP, FN, TN. The right panel plots the resulting ROC curve.

ROC curve construction from a ranked list of 20 emails. Each row sets the threshold at that score; the curve traces TPR vs. FPR as the threshold sweeps from high to low. The highlighted point (T = 0.54) has the highest accuracy but is not necessarily the best operating point.

Figure 3:ROC curve construction from a ranked list of 20 emails. Each row sets the threshold at that score; the curve traces TPR vs. FPR as the threshold sweeps from high to low. The highlighted point (T = 0.54) has the highest accuracy but is not necessarily the best operating point.

Key Landmarks

PointMeaning
(0,0)(0, 0)Threshold so high that nothing is predicted positive (always predict negative)
(1,1)(1, 1)Threshold so low that everything is predicted positive (always predict positive)
(0,1)(0, 1)Perfect classifier: all positives detected, no false positives
DiagonalRandom classifier: TPR = FPR at every threshold (no discrimination)

A curve that hugs the upper-left corner represents a strong classifier: it achieves high TPR before FPR rises appreciably. A curve that follows the diagonal adds no information beyond random guessing.

Reading the ROC Space

Four classifiers (A-D) mapped from their confusion matrices to ROC space. The upper-left region represents high sensitivity with low false-positive rate; the diagonal represents random performance.

Figure 4:Four classifiers (A-D) mapped from their confusion matrices to ROC space. The upper-left region represents high sensitivity with low false-positive rate; the diagonal represents random performance.

Figure 4 places four classifiers at different operating points:

The biometric face-unlock from Evaluating Text Classifiers would appear as a point near the left edge of ROC space: FPR=0.2%FPR = 0.2\%, TPR=98%TPR = 98\%. The requirement “FPR below 0.5%” constrains the operating point to a narrow vertical strip on the left side of the plot.

AUC: Area Under the ROC Curve

The ROC curve shows performance across all thresholds. Summarizing it into a single number gives the Area Under the Curve (AUC).

AUCInterpretation
1.0Perfect separation: every positive scores higher than every negative
0.5Random: scores carry no discriminative information
< 0.5Worse than random (flip predictions to improve)

AUC is useful for comparing classifiers without committing to a threshold: a system with higher AUC has better score separation overall. However, AUC does not say which operating point the system will use in production, and two classifiers with the same AUC can have very different curves (one may be better at low FPR, the other at high TPR).

Choosing an Operating Point

The ROC curve shows what is possible. Selecting one point on it requires a decision criterion from outside the mathematics:

Maximize accuracy. Compute (TP+TN)/(P+N)(TP + TN) / (P + N) at each threshold and pick the maximum. This works when classes are balanced and error costs are symmetric. In the spam example from Figure 3, threshold 0.54 gives TP=5TP = 5, TN=9TN = 9, accuracy =14/20=70%= 14/20 = 70\%, the highest value in the table. But as we discussed, the spam filter should rather optimize for a low FPR while maintaining acceptable TPR, not simply maximize accuracy.

Constrained optimization. Fix one rate and optimize the other. The face-unlock requirement “FPR≤0.2%FPR \leq 0.2\%” defines a vertical line on the ROC plot; the operating point is the highest TPR value at or to the left of that line. This is the standard approach when one error type has a hard cost limit.

Youden’s J-statistic. Maximize J=TPR−FPRJ = TPR - FPR, which is the vertical distance from the diagonal (random classifier). The point with the highest JJ is geometrically closest to the upper-left corner and balances sensitivity against specificity without committing to a cost model.

Equal error rate (EER). Find the threshold where FPR=FNRFPR = FNR (equivalently, FPR=1−TPRFPR = 1 - TPR). This is the point where the ROC curve crosses the anti-diagonal from (0,1)(0, 1) to (1,0)(1, 0). EER is commonly reported in biometric and speaker-verification systems as a single-number summary that does not require choosing which error is worse.

In all cases, threshold selection should be performed on a validation set separate from the test set used to report final metrics. Optimizing the threshold on the test set inflates reported performance.