Nitsor journal. nitsor.com/blog/dataset-versioning-ct-mri. 27 Sep 2026.

Nitsor journal

Medical imaging dataset versioning: reproduce the scans and labels behind a training run

Record the exact scans and labels each training run reads. This public-data example follows the original inputs through later label and class changes.

RSS feed
Three public scans side by side: the airframe's pressure bottle, the lung nodule and an apple. Each carries a model's proposal dashed in orange, and the first two also carry their published outlines in white.
Real scans: Gruber et al., Fraunhofer EZRT (CC BY 4.0 (opens in a new tab)), Armato et al., "Data From LIDC-IDRI", TCIA (CC BY 3.0 (opens in a new tab)), Schut et al., CWI and GREEFA (CC BY 4.0 (opens in a new tab)), modified. White: published labels. Dashed: SAM 2.1 proposals, run by Nitsor outside the product. Sources

Define what a training run must recover

Record which scans and label revisions a training run read so another engineer can repeat the experiment.

Dataset version, in plain words

A fixed selection of labelled scans that a training job keeps using while labelling continues.

A folder name rarely gives the whole answer. The files may still be there, but since the run someone may have corrected a mask, changed a class definition or replaced a source file. A saved query is no better: it returns different members once new scans arrive.

To repeat the data read, an engineer needs the source scans, exact label revisions and class definitions. Record the members of the training, validation and test selections, along with the selection rule, such as accepted work only.

The selection matters as much as the masks. Two experiments can use identical labels on different cases. If the record names only a project, that difference is invisible when someone compares the results.

Fix a selection, then change the labels

The worked example uses two public scans. From LIDC-IDRI-0003, a chest CT, it takes slice 64 and the four radiologists' published outlines of one nodule, each kept as its own label. From the Me 163 V5 industrial scan it takes slice 300 and the published outline of the pressure bottle, instance 138. Figure 1 shows both, with a third sample, as one history.

In this history, each scan has a commit. Each model proposal has its own unmerged branch, so no proposal is a member. The tag marks a version with the seven members listed in Table 1. In Nitsor, creating a Dataset Version freezes the selection at one commit into a manifest. A training run, call it run A, pins that version by its ID and manifest hash.

What a fingerprint is

The first eight characters of a SHA-256 hash of the file. These shortened fingerprints identify files in this example; use the full digest to verify their bytes.1

Table 2 lists three changes made after the version was fixed. After all three, run A still reads exactly the seven members it read on the first day.

A diagram of the label history of three scans, described in the caption below.

Figure 1. The worked example as one history

Three public scans: the airframe section, the chest CT and an apple. Each scan is a commit. Each model proposal, drawn dashed (SAM 2.1, run outside Nitsor), sits on its own branch and is never merged. Each published label is its own branch: the four radiologists keep four branches, and the apple, whose dataset publishes a browning score rather than an outline, carries the score. The labels join the dataset line, the tag dv-sample-01 fixes the version, and a training input preview reads the tag. This guide's seven members are the chest CT and airframe rows.

Hashes are sha256 of the files shown. Data: Gruber et al., Fraunhofer EZRT (CC BY 4.0 (opens in a new tab)), Armato et al., "Data From LIDC-IDRI", TCIA (CC BY 3.0 (opens in a new tab)), Schut et al., CWI and GREEFA (CC BY 4.0 (opens in a new tab)), modified. Sources

Table 1The members of version dv-sample-01. A simplified manifest, not the API schema.
MemberLabelSourceFingerprint
LIDC-IDRI-0003, slice 64none, the scan itselfChest CT, TCIA8022df54
LIDC-IDRI-0003, slice 64reader 1 outlinePublished annotation8d7ed8a4
LIDC-IDRI-0003, slice 64reader 2 outlinePublished annotation1475d738
LIDC-IDRI-0003, slice 64reader 3 outlinePublished annotation0a5d92cb
LIDC-IDRI-0003, slice 64reader 4 outlinePublished annotation1a0f2e8e
Me 163 V5, slice 300none, the scan itselfIndustrial CT, Fraunhofer EZRT48433d90
Me 163 V5, slice 300instance 138, pressure bottlePublished annotation5d6201fd

Version labels and class definitions together

The first change is a correction. A reviewer, M. Chen in this example, tightens the edge of reader 4's outline and saves it. That is a new revision with M. Chen as its author, and the original stays in the history. Run A keeps reading the original, because the version names that revision and not the label in general.

Commit, in plain words

A saved revision, with its author, its time and the revision before it.

The second change is a definition. Suppose an industrial team splits porosity into gas porosity and shrinkage. The team must decide how existing porosity labels map to the new classes; changing the class list alone does not make that decision. Keep the earlier definitions so you can identify the ones run A trained on. In Nitsor, binding a taxonomy writes a checkpoint commit, and annotation commits hash the taxonomy and workflow versions in force.

The third change is a new scan. A saved query can pick it up if it matches the selection rule. Run A's fixed version keeps its original members. A new version includes the scan only if the rule selects it. Record each version's member list to compare the inputs. Dataset versioning describes how to fix the inputs for each run.

Table 2Three changes after the version was fixed.
Change on the branchRun A readsA new version reads
M. Chen, reviewing, corrects the edge of the reader 4 outlineThe original reader 4 outline, 1a0f2e8eThe corrected revision, with reader 4 and M. Chen both named on it
The class list splits one class into twoThe class list it was fixed withThe new class list; existing labels change only if the team remaps them
A new scan joins the branchThe seven members above, and nothing elseThe new selection, if the rule includes the scan

Pin the training job and check what it reads

A training job should name its input by Dataset Version ID and manifest hash. In Nitsor, a Dataset Version freezes a selection at one commit into a manifest with a SHA-256 digest per 16 MiB chunk. A read carries the manifest hash, and a reader can check every chunk against the manifest before any byte reaches a decoder, cache or tensor.2

What a matching hash proves

A matching full digest verifies the bytes against the version's manifest. Reviewers still need to judge whether the labels are correct.

Store the version ID and its manifest hash with the model artefact. When someone later asks which labels trained the model, the answer is those two values. When two runs disagree, compare their versions first. If the versions match, the difference is somewhere else.

Record what happens outside the dataset

A fixed dataset version records the annotation side of the experiment. It does not record what your pipeline does after the read. Resampling, intensity windowing, augmentation, patch sampling and the training code itself each need their own recorded configuration.

The LIDC example shows why. Its slices are 2.5 mm apart and its pixels 0.820 mm across. A pipeline that resamples to 1 mm cubes produces different training arrays from the same version. Keep the resampling settings with the run, beside the version ID.

Decide how the reader outlines become a target, and record that too. Four separate masks, a combined mask by a stated rule, or one reader chosen as the reference are three different training targets from the same seven members. CT annotation quality control works through that choice on these outlines.

A reproducibility checklist

Before a training run starts, record each of the following so another engineer can repeat the data read.

  1. The dataset version ID and manifest hash the run pins.
  2. The selection rule, such as accepted work only.
  3. The members of each split, and how specimens or patients were kept apart.
  4. The class list the labels were drawn under.
  5. How several readers' labels became one training target, if they did.
  6. The preprocessing configuration: resampling, windowing, augmentation.
  7. The training code revision.
  8. Where the model artefact records all of the above.

If your team needs to trace a model back to its training data, we can walk through a fixed version and a label change on a public scan. Request a walkthrough.

In the product: Dataset versioningRead the dataset guide (opens in a new tab)

Notes and sources

  1. Fingerprints are the first eight hex characters of the SHA-256 of each file shown: the model-input image for a scan, and the published label array on the slice shown for a label. Armato et al., "Data From LIDC-IDRI", TCIA, doi:10.7937/K9/TCIA.2015.LO9QL9SX (opens in a new tab), CC BY 3.0 (opens in a new tab). Gruber et al., Fraunhofer EZRT Me 163 V5, doi:10.5281/zenodo.10651746 (opens in a new tab), CC BY 4.0 (opens in a new tab). Back to text
  2. The dataset release guide (opens in a new tab) describes how a version is made, pinned and read. Back to text

Read next

Pictures: real clinical CT, Armato et al., "Data From LIDC-IDRI", TCIA, CC BY 3.0 (opens in a new tab), modified; real industrial CT, Gruber et al., Fraunhofer EZRT, CC BY 4.0 (opens in a new tab), modified. Sources

Show us what you inspect.

Start with a public CT or MRI scan. See how a model suggestion becomes a reviewed label and a fixed training dataset.