A 344 KB linear probe on frozen features, ahead of a fine-tuned ResNet-50

Linear(5376, 14) fitted on frozen google/gemma-4-31B-it hidden states. The backbone is untouched: no fine-tuning, no radiology pretraining, no augmentation. All 112,120 images of NIH ChestX-ray14, official test_list.txt.

ChestX-ray14, official split mean AUROC trainable params
Wang et al. 2017 (dataset authors) 0.7451 ~25 M, fine-tuned end to end
this probe 0.7590 75,278
view-position only 0.5896 shortcut baseline
shuffled labels 0.5002 refit floor

Ahead on 12 of 14 findings. 14 of 14 clear the view-position baseline.

The whole result refits from the published states in 48 seconds on a MacBook CPU. No GPU is needed to reproduce it, only to encode the images once.

Read the split before comparing

CheXNet's 0.8414 and Yao's 0.8027 are on a different test set. CheXNet section 5: "We randomly split the dataset into training (70%), validation (10%), and test (20%) sets." That is their own random partition, not the official list.

Which split is harder is not established. Wang scored 0.7381 on a random partition and 0.7451 on the official one, so for their model the official split was marginally easier. The only supportable statement is that the two are not comparable, so exactly one split-matched row appears above.

Do not compare 0.7590 to 0.8414.

Per-finding

finding AUROC 95% CI shuffled view-only n_pos
Emphysema 0.8650 [0.849, 0.880] 0.4340 0.5740 1093
Pneumothorax 0.8465 [0.832, 0.859] 0.5252 0.5870 2665
Cardiomegaly 0.8221 [0.798, 0.844] 0.5347 0.5185 1069
Edema 0.8170 [0.798, 0.837] 0.4975 0.7011 925
Effusion 0.7849 [0.773, 0.796] 0.4740 0.5274 4658
Hernia 0.7828 [0.692, 0.869] 0.5174 0.6564 86
Fibrosis 0.7538 [0.726, 0.781] 0.4836 0.6310 435
Mass 0.7423 [0.718, 0.766] 0.5159 0.5502 1748
Pleural_Thickening 0.7347 [0.714, 0.754] 0.5031 0.5920 1143
Atelectasis 0.7248 [0.710, 0.738] 0.4901 0.5140 3279
Consolidation 0.7107 [0.695, 0.727] 0.5032 0.6377 1815
Nodule 0.6956 [0.674, 0.715] 0.5137 0.5768 1623
Infiltration 0.6862 [0.674, 0.696] 0.5048 0.6034 6112
Pneumonia 0.6600 [0.637, 0.686] 0.5052 0.5853 555

Intervals are a patient-level cluster bootstrap. The test split is 25,596 films from 2,797 patients, about 9 each, and those films are not independent. Resampling images instead of patients gives intervals roughly 1.5x too narrow.

The view-only baseline is folded (max(vw, 1-vw)). Hernia's raw view-only AUROC is 0.3436, which is 0.6564 of shortcut once flipped, and reporting the raw figure would have flattered the probe.

Usage

The probe expects a mean-pooled hidden state from the same backbone and layer, and the checkpoint carries the train-split mu and sd it needs. Applied to any other features, or without that normalisation, the scores mean nothing.

import torch
from huggingface_hub import hf_hub_download

ck = torch.load(hf_hub_download("RiverRider/srt-cxr14-linear-probe",
                                "cxr14_probe.pt"), weights_only=True)
x = (state - ck["mu"]) / ck["sd"]          # state: (n, 5376) from gemma-4-31B-it
p = torch.sigmoid(x @ ck["W"] + ck["b"])   # (n, 14), order in ck["findings"]

Precomputed states for all 112,120 images, the manifest with the official split, and the fitting script are in RiverRider/srt-cxr14-frozen-probe.

Scope, and what this is not

Detection, not early detection. These labels describe what is visible in the image in front of you. Nothing here speaks to catching disease before it is apparent, which needs longitudinal data with outcomes.

Not a diagnostic device. Research artifact. No clinical validation, no prospective evaluation, no regulatory clearance.

Labels are NLP-mined from radiology reports by the dataset authors. Every model on this benchmark inherits that ceiling.

One backbone. The probe is deliberately linear, because anything stronger starts measuring the probe rather than the representation.

Banked negatives

hypothesis result
Attention-style pooling beats mean for focal findings Falsified. Focal mean falls 0.0537 under max-pool and 0.0225 under top16, at every depth tested
Readout depth matters No. 0.7600 to 0.7605 across 0.4 / 0.6 / 0.8 of backbone depth

Part of the SRT program, https://github.com/space-bacon/SRT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RiverRider/srt-cxr14-linear-probe

Finetuned
(244)
this model

Dataset used to train RiverRider/srt-cxr14-linear-probe

Space using RiverRider/srt-cxr14-linear-probe 1

Collection including RiverRider/srt-cxr14-linear-probe

Evaluation results

  • Mean AUROC over 14 findings on NIH ChestX-ray14 (official test_list.txt)
    self-reported
    0.759