Review the slices a model is unsure about, not all of them.

Your expert reads ninety slices instead of nine hundred, and the ninety are chosen by a rule, not a hunch. The cut is measured on cases the model has not seen, so it comes with a number you can hold us to. Ask six months later why nobody looked at slice 412 and the answer sits on the record, next to the labels.

What it decides
which slices a person reviews
How the cut is set
measured on held-out cases
What it is not
the model's own confidence score
What gets written down
the error budget, the sample unit, the model and data versions
calibration
Checking a model's confidence against what actually happened, on cases it was not trained on. A model that says ninety per cent should turn out right about ninety times in a hundred. Most are not.
distribution-free risk control
A way of setting the cut so the error rate stays under a number you choose. It is measured on held-out cases and assumes nothing about the shape of your data. It writes down what it counted and what it assumed.

Move the cut. The queue and the record move with it.

The bars are how unsure a model was, slice by slice, across one 512-slice series. Everything to the right of the cut goes to a person. Everything to the left the model handles alone. Drag the cut and the queue, the answer about slice 412 and the record underneath are all recomputed.

score histogram 512 sliceserror budget 5.0 per cent

A person looks at 90 of 512 slices. The model covers the other 422.

Why did nobody look at slice 412? Its score was 0.31, below the cut at 0.40, so the model handled it, inside an error budget of 5.0 per cent.

method
distribution-free risk control
sample unit
slice
assumption
exchangeable within series
error budget at this cut
5.0 per cent
routed to a person
90
covered by the cut
422
model
seg 0.4.1
data commit
sha256:137693998ca4
cut
0.40
Synthetic sampleThe score distribution and the model and data versions are made up for this page. The counts are arithmetic on 512 slices.

The same 512 slices, two rules for who looks.

Drag the line, or click anywhere in the queue. Left of the line the queue is picked by the model's own confidence score. Right of it by a cut measured on held-out cases. Same slices, same series, two rules. The cells that change colour as the line passes are the slices the two rules disagree about.

one series 512 sliceseach cell is one slice

left the model's own scorea measured cut right

The model's own score, above 0.75

27 of 512 to a person

Fewer slices, and no number behind them. A confidence score is what the model says about itself. Nothing has checked whether it is right.

A cut measured on held-out cases

90 of 512 to a person

More slices, and a promise attached. At most five errors in a hundred on cases the model has not seen, with the sample unit and the assumption written down beside it.

Synthetic sampleThe score distribution is made up for this page. The two counts are arithmetic on the same 512 slices.

Before we measure anything, we agree what counts as an error.

A budget of five in a hundred is meaningless until somebody says five of what. That conversation happens first, with you, and the answer goes on the record with the numbers.

one errorWhat a wrong slice is on your partsyou define
weightingWhether a missed defect and a false one weigh the sameyou decide
unitA slice, a part, a series or one defect classyou choose
the casesWhich held-out scans the measurement runs onyours

The figure above counts a slice as the unit. On your data it might be a part, a series or a single defect class, and that choice changes the number. So it is settled first, in writing, and it goes on the record beside the measurement.

The guarantee that comes out is an average across the slices the cut covered. It is not a promise about any one part, and no honest method gives you that. Say that sentence to your quality manager before you set a budget, because it is the sentence they will ask you about.

The routing decision is stored per slice, next to the labels, in the same history. So the question is never "what was the policy in March". It is "what happened to slice 412 of this part", and the answer comes back for that part.

What you get

  • A queue with a reason

    Every slice a person sees is in the queue because of a rule, not a hunch. The rule is stored, not remembered.

  • An answer about the ones you skipped

    The record names the target risk, the sample unit, the assumption, the model version and the data version the cut was measured against.

  • A cut you can move

    Raise the target risk and fewer slices reach a person. Lower it and more do. The trade is yours to make and it is visible.

  • A record that outlives the model

    The routing decision sits next to the labels in the same history, so a later reader can check it against the release they were given.

Where these numbers come from

Every score in the two figures above comes from a made-up distribution. We have not run a model on this airframe and we are not going to pretend we have.

The counts are real arithmetic on 512 slices, and the series is real: Fraunhofer EZRT XXL-CT Me 163, subvolume V5, published under CC BY 4.0.

A cut measured on somebody else's data tells you nothing about yours, which is why we set this up with you, on your scans, rather than shipping a number.

Bring one set of scans. Leave with a record.

Thirty minutes on your own scans, on our hosted instance, no slides. If we are not a fit, we say so on the call.