Medical imaging dataset versioning: reproduce the scans and labels behind a training run
Record the exact scans and labels each training run reads. This public-data example follows the original inputs through later label and class changes.

Define what a training run must recover
Record which scans and label revisions a training run read so another engineer can repeat the experiment.
A fixed selection of labelled scans that a training job keeps using while labelling continues.
A folder name rarely gives the whole answer. The files may still be there, but since the run someone may have corrected a mask, changed a class definition or replaced a source file. A saved query is no better: it returns different members once new scans arrive.
To repeat the data read, an engineer needs the source scans, exact label revisions and class definitions. Record the members of the training, validation and test selections, along with the selection rule, such as accepted work only.
The selection matters as much as the masks. Two experiments can use identical labels on different cases. If the record names only a project, that difference is invisible when someone compares the results.
Fix a selection, then change the labels
The worked example uses two public scans. From LIDC-IDRI-0003, a chest CT, it takes slice 64 and the four radiologists' published outlines of one nodule, each kept as its own label. From the Me 163 V5 industrial scan it takes slice 300 and the published outline of the pressure bottle, instance 138. Figure 1 shows both, with a third sample, as one history.
In this history, each scan has a commit. Each model proposal has its own unmerged branch, so no proposal is a member. The tag marks a version with the seven members listed in Table 1. In Nitsor, creating a Dataset Version freezes the selection at one commit into a manifest. A training run, call it run A, pins that version by its ID and manifest hash.
The first eight characters of a SHA-256 hash of the file. These shortened fingerprints identify files in this example; use the full digest to verify their bytes.1
Table 2 lists three changes made after the version was fixed. After all three, run A still reads exactly the seven members it read on the first day.

Figure 1. The worked example as one history
Three public scans: the airframe section, the chest CT and an apple. Each scan is a commit. Each model proposal, drawn dashed (SAM 2.1, run outside Nitsor), sits on its own branch and is never merged. Each published label is its own branch: the four radiologists keep four branches, and the apple, whose dataset publishes a browning score rather than an outline, carries the score. The labels join the dataset line, the tag dv-sample-01 fixes the version, and a training input preview reads the tag. This guide's seven members are the chest CT and airframe rows.
Hashes are sha256 of the files shown. Data: Gruber et al., Fraunhofer EZRT (CC BY 4.0 (opens in a new tab)), Armato et al., "Data From LIDC-IDRI", TCIA (CC BY 3.0 (opens in a new tab)), Schut et al., CWI and GREEFA (CC BY 4.0 (opens in a new tab)), modified. Sources
| Member | Label | Source | Fingerprint |
|---|---|---|---|
| LIDC-IDRI-0003, slice 64 | none, the scan itself | Chest CT, TCIA | 8022df54 |
| LIDC-IDRI-0003, slice 64 | reader 1 outline | Published annotation | 8d7ed8a4 |
| LIDC-IDRI-0003, slice 64 | reader 2 outline | Published annotation | 1475d738 |
| LIDC-IDRI-0003, slice 64 | reader 3 outline | Published annotation | 0a5d92cb |
| LIDC-IDRI-0003, slice 64 | reader 4 outline | Published annotation | 1a0f2e8e |
| Me 163 V5, slice 300 | none, the scan itself | Industrial CT, Fraunhofer EZRT | 48433d90 |
| Me 163 V5, slice 300 | instance 138, pressure bottle | Published annotation | 5d6201fd |
Version labels and class definitions together
The first change is a correction. A reviewer, M. Chen in this example, tightens the edge of reader 4's outline and saves it. That is a new revision with M. Chen as its author, and the original stays in the history. Run A keeps reading the original, because the version names that revision and not the label in general.
A saved revision, with its author, its time and the revision before it.
The second change is a definition. Suppose an industrial team splits porosity into gas porosity and shrinkage. The team must decide how existing porosity labels map to the new classes; changing the class list alone does not make that decision. Keep the earlier definitions so you can identify the ones run A trained on. In Nitsor, binding a taxonomy writes a checkpoint commit, and annotation commits hash the taxonomy and workflow versions in force.
The third change is a new scan. A saved query can pick it up if it matches the selection rule. Run A's fixed version keeps its original members. A new version includes the scan only if the rule selects it. Record each version's member list to compare the inputs. Dataset versioning describes how to fix the inputs for each run.
| Change on the branch | Run A reads | A new version reads |
|---|---|---|
| M. Chen, reviewing, corrects the edge of the reader 4 outline | The original reader 4 outline, 1a0f2e8e | The corrected revision, with reader 4 and M. Chen both named on it |
| The class list splits one class into two | The class list it was fixed with | The new class list; existing labels change only if the team remaps them |
| A new scan joins the branch | The seven members above, and nothing else | The new selection, if the rule includes the scan |
Pin the training job and check what it reads
A training job should name its input by Dataset Version ID and manifest hash. In Nitsor, a Dataset Version freezes a selection at one commit into a manifest with a SHA-256 digest per 16 MiB chunk. A read carries the manifest hash, and a reader can check every chunk against the manifest before any byte reaches a decoder, cache or tensor.2
A matching full digest verifies the bytes against the version's manifest. Reviewers still need to judge whether the labels are correct.
Store the version ID and its manifest hash with the model artefact. When someone later asks which labels trained the model, the answer is those two values. When two runs disagree, compare their versions first. If the versions match, the difference is somewhere else.
Record what happens outside the dataset
A fixed dataset version records the annotation side of the experiment. It does not record what your pipeline does after the read. Resampling, intensity windowing, augmentation, patch sampling and the training code itself each need their own recorded configuration.
The LIDC example shows why. Its slices are 2.5 mm apart and its pixels 0.820 mm across. A pipeline that resamples to 1 mm cubes produces different training arrays from the same version. Keep the resampling settings with the run, beside the version ID.
Decide how the reader outlines become a target, and record that too. Four separate masks, a combined mask by a stated rule, or one reader chosen as the reference are three different training targets from the same seven members. CT annotation quality control works through that choice on these outlines.
A reproducibility checklist
Before a training run starts, record each of the following so another engineer can repeat the data read.
- The dataset version ID and manifest hash the run pins.
- The selection rule, such as accepted work only.
- The members of each split, and how specimens or patients were kept apart.
- The class list the labels were drawn under.
- How several readers' labels became one training target, if they did.
- The preprocessing configuration: resampling, windowing, augmentation.
- The training code revision.
- Where the model artefact records all of the above.
If your team needs to trace a model back to its training data, we can walk through a fixed version and a label change on a public scan. Request a walkthrough.
In the product: Dataset versioningRead the dataset guide (opens in a new tab)
Notes and sources
- Fingerprints are the first eight hex characters of the SHA-256 of each file shown: the model-input image for a scan, and the published label array on the slice shown for a label. Armato et al., "Data From LIDC-IDRI", TCIA, doi:10.7937/K9/TCIA.2015.LO9QL9SX (opens in a new tab), CC BY 3.0 (opens in a new tab). Gruber et al., Fraunhofer EZRT Me 163 V5, doi:10.5281/zenodo.10651746 (opens in a new tab), CC BY 4.0 (opens in a new tab). Back to text
- The dataset release guide (opens in a new tab) describes how a version is made, pinned and read. Back to text

