OrcaDose
Sign in
structuresfree to quote

Reviewing an auto-segmented contour

Dice tells you almost nothing about whether a contour will change the plan. Review where the dose gradient is, not where the volume is.

Built in·updated 2026-08-03

Deep-learning auto-segmentation is fast and uneven, and the usual acceptance numbers do not measure the thing that matters.

The geometric numbers, and their limit. TG-132 gives a common bar — DSC > 0.8 and mean distance to agreement < 3 mm. Useful for saying a model works at all. Nearly useless for a specific patient, because Dice is dominated by the bulk of a structure while dose is decided at its edge. A parotid with excellent Dice and a 4 mm error on the medial surface facing the target is a good contour and a bad plan.

Review in the order dose cares about.

  1. The surface facing the target. Every millimetre there is worth more than the rest of the structure put together.
  2. The first and last slices. Superior–inferior extent is where these models fail most and where a reviewer's eye skips.
  3. Anything the model has probably never seen — post-surgical anatomy, a prosthesis, an unusual position, a structure the training set called something else.

What it is genuinely good at. Reported gains are real: contouring time down by around 30 minutes in a head-and-neck series, and — the finding that matters more — less variation between observers at different institutions. Consistency is where auto-segmentation earns its place, not speed.

Do not skip validation because the vendor did some. Published guidance is explicit that extensive local validation is needed before clinical use; a model trained elsewhere meets your scanner, your protocol and your naming conventions for the first time on your patient.

On this platform. If you contribute a case with auto-generated contours, say so in the description. Everyone who plans it is scored against those contours, and a reviewer who knows they were automated will look at the right places.

Where the numbers come from