Your expert needs to look at ninety of these.

Not all nine hundred, and not a random sample. Calibrated uncertainty picks the ninety, and records why it skipped the rest.

Reviewing every slice costs more expert hours than anyone has. A random sample misses the defects that matter, because defects do not spread themselves evenly. This is the third option, and it comes with a stated risk bound.

Become a design partner
Review priority
  • Bucket 1 of 6, auto-accept candidate: 810 slices
  • Bucket 2 of 6: 14 slices
  • Bucket 3 of 6: 18 slices
  • Bucket 4 of 6: 20 slices
  • Bucket 5 of 6: 18 slices
  • Bucket 6 of 6, expert required: 20 slices
6 buckets, from auto-accept candidate to expert required

One cell is one axial slice. 900 slices, 90 routed. Synthetic sample.

Three things make it honest, and hard to copy.

Calibrated, not confident

Softmax scores are not calibration. Treating them as one is how a team ends up trusting a number that means nothing. Routing runs on distribution-free risk control, and the record names the sample unit it counts and the exchangeability assumption it rests on, that calibration slices and routed slices come from one population.

Refusal is a first-class result

Too few independent samples, or a shift in the data, and the assumptions break. Nitsor then says it cannot claim, instead of printing a number anyway. A tool that always produces a confidence is not being careful. It is being decorative.

Routing is itself evidence

Why did no human look at slice 412? Nitsor records the decision, its basis, the model version, the execution stack and the data commit together, so the question has an answer twelve months later instead of a shrug.

What it costs, and what it does not buy.

On a 900-slice CT scan where calibration clears 810 slices, the expert opens ninety of them. The ninety are the cheap part. The bound on what the other 810 might hide, and the record of how it was set, is what survives an audit.

Routing does not make the model correct. It does not remove the need for a qualified person. And it does not reach past the population the calibration was computed on. Change the scanner or the reconstruction kernel and the bound has to be earned again.

Slices an expert opens, one study

Illustrative
Review every slice900
Uncertainty-routed90
Bars are slice counts, not measured time. This is a drawing of how routing splits a study, not a benchmark result. Synthetic sample.

Routing is what a Certificate unlocks.

The threshold that lets a prediction skip review is not a slider someone drags. It is set in the units of a calibration Certificate bound to that model version, that execution stack and that data commit. No Certificate, no auto-accept. A stronger Certificate widens the band, a weaker one narrows it.