CT annotation quality control: measure reader agreement and review disagreements
Four radiologists outlined one nodule and drew four masks. Measure where they agree, then review where they do not, with a name on every decision.

Keep each reader's annotation
Four radiologists can outline the same finding and produce four different masks. A quality process has to keep those differences long enough for someone to understand them.
The case in this guide is LIDC-IDRI-0003, a public chest CT with published reader outlines. The picture at the top of this guide shows the four outlines of one nodule in white, on the slice the scanner numbered 64. They overlap along much of the edge but differ in some regions. Reader 4 drew the largest outline and marked the nodule on 10 slices. The other three marked it on 7.
A public collection of chest CT scans in which up to four radiologists outlined lung nodules. These outlines come from the final round, after each reader had seen the others' anonymised marks, and they still differ.1
Store an individual mask for every reader before building any combined reference. Keep the reader identifier, the source series, the slice geometry and the version of the instructions each reader followed. Without those, a disagreement score is hard to read and harder to repeat.
A combined mask answers a different question. A common rule for LIDC keeps every pixel that at least two of four readers included. On this slice that gives 479.1 mm², against 411.1 to 580.0 mm² for the individual readers. Record that rule with the combined mask so the next team can reproduce it.
Choose the comparison unit
Decide what you compare before you calculate anything. One slice, one object across all its slices, and a whole volume give three different answers, and a report that mixes them is hard to interpret.
This guide starts with one slice because readers can compare all four outlines by eye. That comparison does not describe agreement across the whole object. One object across its slices is usually the unit that matters for training, since the model learns the whole shape. A whole volume adds empty slices, and those need a rule of their own: two readers who both leave a slice empty agree completely, and averaging that in makes every case look better.
Table 1 already shows why the unit matters. The readers disagree about area on this slice by up to 169 mm², and reader 4 also outlines the nodule on three more slices than the others. A single-slice score cannot see that second disagreement at all.
This series has 0.820 mm pixels and 2.5 mm between slices, so one pixel on the slice is about 0.673 mm².
| Reader | Area on this slice, mm² | Longest diameter, mm | Slices outlined |
|---|---|---|---|
| Reader 1 | 418.6 | 31.4 | 7, InstanceNumber 62 to 68 |
| Reader 2 | 411.1 | 31.2 | 7, InstanceNumber 62 to 68 |
| Reader 3 | 454.2 | 27.6 | 7, InstanceNumber 62 to 68 |
| Reader 4 | 580.0 | 36.1 | 10, InstanceNumber 61 to 70 |
Calculate overlap without hiding disagreement
With four readers there are six pairs. On this slice their IoU runs from 0.71 to 0.85.2 Keep all six beside the masks. A reader can then see which pair produced each value and where the outlines diverge. The average of 0.78 does not identify the pairs that disagree most.
IoU is the overlap of two masks divided by their combined area. Dice counts the overlap twice against the sum of both areas, so it is at least as high for the same pair.
In Table 2, every pair that includes reader 4 has lower overlap than the three pairs among readers 1, 2 and 3. At adjudication, a radiologist should check whether reader 4 included tissue the others left out, then ask whether the difference came from the instructions or from judgement about the edge.
The dashed outline in the picture at the top of this guide is a SAM 2.1 proposal from one box. Against the four readers its IoU is 0.79, 0.80, 0.84 and 0.84, respectively.3
Look at where the pixels agree as well as how much. Of the 880 pixels any reader marked on this slice, 547 were marked by all four, and 168 by only one. Most of those 168 lie within three pixels of the outer edge (Figure 1). Inspect that band to see where the readers disagree.
Boundary measurements
A two-pixel boundary shift changes IoU less on a large mask than on a small nodule. For small findings, add a boundary measure such as the largest distance between two outlines, and report it in millimetres at the recorded spacing.

Figure 1. Reader agreement on InstanceNumber 64
- 1 reader168 pixels
- 2 readers81 pixels
- 3 readers84 pixels
- All 4 readers547 pixels
- Published outline, one per reader
880 pixels were marked by at least one reader. Each square in the picture is one 0.82 mm pixel. The veil is derived from the published outlines; the scan is dimmed so its four steps can be told apart.
Real clinical CT, Armato et al., "Data From LIDC-IDRI", TCIA, CC BY 3.0 (opens in a new tab), modified: cropped, windowed and overlaid by Nitsor. Sources
| Readers | IoU | Dice |
|---|---|---|
| 1 and 2 | 0.85 | 0.92 |
| 2 and 3 | 0.82 | 0.90 |
| 1 and 3 | 0.81 | 0.89 |
| 3 and 4 | 0.75 | 0.86 |
| 1 and 4 | 0.71 | 0.83 |
| 2 and 4 | 0.71 | 0.83 |
Decide how a disagreement gets resolved
Write down how adjudication resolves a disagreement before the first study arrives. Overlap scores identify differences but cannot determine which outline to use.
Your written procedure could define four outcomes, separate from Nitsor's verdicts. Accept one reader's outline as the reference. Correct it into a new outline and record the reason. Return the case to the original reader with a note about the instruction. Or keep the disagreement as data, because some findings have no single agreed edge. An evaluation should account for that uncertainty.
Decide which cases need review. Your procedure might require review below an agreement threshold and a random sample of the remaining cases. Sampling can reveal errors even where readers agree. The agreement calculations in this guide ran outside Nitsor. Routed review applies the same rule to model output: a threshold and a sampling rate you set, with the rule kept on each task.
Record which revision the reviewer judged
A review record should separate the original reader annotation from any later correction. Keep both, name the person who made the change, and attach the verdict to the revision that was inspected.
A saved version of a label, with its author, its time and the version before it.
For a case like reader 4's outline above, send the study to a consensus read. Two to sixteen readers label it, blind to each other by default, and Nitsor scores their agreement: Dice or IoU for masks, distance in millimetres for points and exact match for each field.
Your team sets a criterion for each metric, a minimum overlap or a maximum distance, whether readers work independently or check and correct, who may adjudicate and what each reader sees. When reads miss a criterion, the study goes to an adjudicator. The adjudicator takes one reader's labels, keeps the original or edits their own, and History records the decision with a name and a time.
A single second read is recorded the same way. Automatic routing sends the work to someone other than its author; if an operator names the author instead, the record shows it was a self-review. The reviewer records accept, correct or reject with a name and a time. A disagreement your team keeps as data carries no verdict; write that decision into the task instead.
At a review stage, a reviewer records accept, reject or correct per annotation row, bound to the annotation hash judged. The reviewer sees who authored each annotation, and finishing the review records a commit under the reviewer's name. Editing an annotation changes its hash, so an earlier verdict stops applying; the verdict remains in the history. The agreement numbers in this guide were calculated from the published outlines with a short script outside Nitsor. Keep those outlines and the calculation inputs to trace each result.
An annotation QA worksheet
Fill this in once per project, before the first comparison. The answers belong in the task instructions, where every reader and reviewer can see them.
- Name the finding or structure, and link the instruction version readers follow.
- Keep one mask per reader, with the reader, series and slice geometry.
- Choose the comparison unit: one slice, one object across its slices, or a volume.
- Say how empty slices count before averaging anything.
- Report every pairwise score, not only the mean.
- Add a boundary measure in millimetres for small findings.
- State the rule behind any combined reference, such as at least two of four readers.
- List the four procedure outcomes above and who owns each.
- Set the agreement threshold for review and a sampling rate for the rest.
- Attach every verdict to the revision the reviewer saw.
If your team reviews CT or MRI labels with more than one reader, we can walk through this worksheet on a public case close to your data. Request a walkthrough.
In the product: Routed reviewRead the annotation guide (opens in a new tab)
Notes and sources
- Armato et al., "Data From LIDC-IDRI", The Cancer Imaging Archive, doi:10.7937/K9/TCIA.2015.LO9QL9SX (opens in a new tab), CC BY 3.0 (opens in a new tab). Case LIDC-IDRI-0003: 140 slices, 0.820 mm pixels, 2.5 mm between slices. The four outlines are the readers' published annotations of one nodule. Windowed and rendered by Nitsor. Back to text
- Calculated by Nitsor from the published outlines, filled on the 512 by 512 grid of InstanceNumber 64, on 27 September 2026. IoU and Dice are rounded to two places. The agreement counts on this slice are 168, 81, 84 and 547 pixels for one, two, three and four readers. Areas and diameters in Table 1 are derived from the same outlines. Back to text
- SAM 2.1 hiera-tiny from Meta, Apache 2.0 (opens in a new tab), run by Nitsor outside the product on this slice with one box drawn from the readers' combined extent plus 3 mm. One run on one slice. It is an example of comparing a proposal against several references, not a measure of model accuracy. Back to text

