Geometric accuracy and time efficiency evaluation of a deep learning auto-segmentation system for head and neck radiotherapy: a single-center validation study
Highlight box
Key findings
• The prototype deep learning contouring (DLC) system achieved robust segmentation for large, high-contrast organs [eyes, brainstem, parotid glands; Dice similarity coefficient (DSC) ≥0.80] but suboptimal performance for some structures [optic chiasm, DSC = 0.44; pharyngeal constrictor muscle (PCM), DSC = 0.42; mandible, DSC = 0.41]. The human-in-the-loop workflow reduced total contouring time by 30% across all organs at risk (OARs) (35.3 → 24.7 minutes) and by 48.6% when restricted to the structures covered by the DLC model (21.8 → 11.2 minutes; both P<0.001). Notably, correlation between objective geometric metrics and subjective clinical acceptability was only modest (|τ| ≤0.42), with only 4 of 28 correlations surviving false discovery rate correction.
What is known and what is new?
• Deep learning-based auto-segmentation has shown promise in reducing contouring time for radiotherapy planning, but most validations focus on geometric accuracy alone without evaluating the discordance between objective metrics and clinical judgment.
• This study evaluates each stage of the contouring workflow, unmodified DLC, manually modified DLC, and de novo manual contours, in a clinically challenging head and neck cancer cohort (64% stage III–IV), providing evidence that geometric metrics such as DSC and Hausdorff distance alone are insufficient to predict clinical acceptability.
What is the implication, and what should change now?
• DLC systems should be implemented as supportive tools within a human-in-the-loop workflow rather than as replacements for expert delineation, particularly for small, critical structures. Future validation studies should incorporate dosimetric evaluation and multi-dimensional assessment frameworks beyond geometric metrics alone to more accurately capture clinical utility.
Introduction
Globally, there are approximately 900,000 newly diagnosed head and neck cancer (HNC) cases and an estimated 400,000 deaths annually (1). High-quality radiotherapy (RT) is essential for optimal treatment outcomes. Several advanced RT techniques have been developed, including intensity-modulated RT, volumetric modulated arc therapy, and particle therapy. Accurately identifying and segmenting organs at risk (OARs) is pivotal for enhancing the efficacy and safety of sophisticated modern RT approaches. The association between OARs and a broad spectrum of RT-related side effects is extensively documented, such as dysphagia, dependence on feeding tubes, xerostomia, osteoradionecrosis, dental complications, and fatigue (2-6). Correct segmentation of OARs necessitates a thorough understanding of both basic and complex postoperative anatomy. As a result, manual contouring (MC) is time-consuming and demands considerable attention from practitioners (7,8). Significant inter-observer and intra-observer variability further complicates the development of a reliable dose-response relationship for clinical application (9-11). Moreover, the rapid progression in daily adaptive RT significantly increases the need for precise OAR segmentation. Consequently, there is a compelling need to validate automated tools that can address these efficiency bottlenecks without compromising safety.
Various strategies have been developed to enhance the accuracy of OAR segmentation in RT planning, including atlas-based (12-14) and deep learning-based methods (15,16). Atlas-based auto-segmentation tools, widely adopted in clinical settings, utilize a pre-established image library to register and adapt expert-contoured OARs to new patient images. However, these tools often yield unsatisfactory results, particularly for patients with significantly altered anatomy due to extensive surgeries or advanced tumor invasions (14). In contrast, U-Net-based neural networks (17) represent the state-of-the-art in medical image segmentation. Demonstrating practicality and reliability, deep learning contours (DLC) have been effectively implemented for various anatomical sites, including the chest (18), abdomen (19,20), and pelvis (21,22). At our institution, a prototype DLC system, developed by GE Healthcare and based on the U-Net architecture, has been integrated into the clinical workflow for HNC patients as part of a research initiative. This integration aims to evaluate the system’s performance in segmenting OARs and its time efficiency in resource utilization.
While the “human-in-the-loop” workflow, where clinicians review and modify AI-generated contours, is already the established standard of care, the specific performance characteristics of these commercial prototypes on complex, real-world HNC anatomies remain underexplored. Therefore, rather than evaluating only the final clinical output, this study separately evaluates each stage of the contouring workflow to characterize where and to what extent manual intervention is required. We evaluate and compare three contouring strategies: unmodified deep learning contouring (DLC-u), deep learning contouring with manual modification (DLC-m), and de novo MC. By comparing DLC-u with DLC-m, we can quantify the specific degree of manual intervention required to achieve clinical safety, while the comparison against MC establishes the true time-efficiency gained. Furthermore, because strict geometric accuracy does not always perfectly translate to clinical acceptability, a primary objective of this study is to investigate the correlation between objective geometric errors [measured by Dice similarity coefficient (DSC) and Hausdorff distance (HD)] and subjective clinical usability (assessed via a 5-point Likert scale). This dual approach provides critical insights into the real-world readiness and limitations of deep learning auto-segmentation in complex head and neck RT.
Methods
Patients
Between September 2019 and January 2021, 50 consecutive patients undergoing intensity-modulated proton therapy (IMPT) for HNC were enrolled at Linkou Chang Gung Memorial Hospital, Taoyuan, Taiwan. Eligible patients met all of the following inclusion criteria: (I) age ≥18 years; (II) pathologically confirmed head-and-neck malignancy; (III) treated with IMPT in either the definitive or postoperative setting; and (IV) availability of a complete planning computed tomography (CT) scan. Patients were excluded if they (I) underwent re-irradiation to overlapping head-and-neck volumes; (II) had non-malignant lesions; or (III) were treated with palliative intent. The primary tumor sites included the nasopharynx, oropharynx, oral cavity, hypopharynx, larynx, and cases of unknown primary origin. Detailed clinical characteristics of the cohort are summarized in Table 1. T3–T4 disease was present in 46% of patients, and 42% had N2 or higher nodal involvement. Overall, 64% had stage III–IV disease [American Joint Committee on Cancer (AJCC) 8th edition], and 14% had undergone surgery prior to RT. Pathological staging was used for postoperative patients when available. For each patient, planning CT scans were acquired using the Discovery CT590 RT scanner (GE Healthcare, Chicago, IL, USA) under standardized parameters: 120 kVp, 300 mAs, and a voxel size of 0.98×0.98×1.25 mm. The scans incorporated the built-in metal artifact reduction (MAR) algorithm for OAR contouring in patients with metal implants in the head and neck region, in accordance with institutional protocol. Contrast enhancement was not allowed.
Table 1
| Variable | Category | No. (%) |
|---|---|---|
| Primary tumor site | Nasopharynx | 20 (40) |
| Oropharynx | 14 (28) | |
| Oral cavity | 10 (20) | |
| Hypopharynx/larynx | 5 (10) | |
| Unknown primary | 1 (2) | |
| T stage† | T0/Tx | 2 (4) |
| T1 | 12 (24) | |
| T2 | 13 (26) | |
| T3 | 9 (18) | |
| T4 | 14 (28) | |
| N stage† | N0 | 16 (32) |
| N1 | 13 (26) | |
| N2 | 12 (24) | |
| N3 | 9 (18) | |
| Stage (AJCC 8th edition) | I | 9 (18) |
| II | 9 (18) | |
| III | 12 (24) | |
| IV | 20 (40) | |
| Treatment strategy | Definitive | 43 (86) |
| Postoperative | 7 (14) |
†, pathological staging was used when available (n=7); clinical staging was used otherwise. AJCC, American Joint Committee on Cancer.
MC
OAR segmentation was conducted by three board-certified radiation oncologists, two of whom specialize in the treatment of HNC. No study-specific training session was conducted prior to the study period, and the participating radiation oncologists relied on their routine clinical expertise and the institution’s existing head-and-neck OAR contouring practice. Each patient’s OAR set was contoured by one of the three oncologists and served as the reference standard. No formal multi-observer consensus mechanism (e.g., STAPLE) was employed. To ensure relevance to real-world clinical practice, the selection of OARs was guided by the NRG-HN001 (NCT02135042) and NRG-HN002 (23) study protocols. The selected OARs included the brachial plexus, brainstem, cochlea, eye, larynx, lens, mandible, optic chiasm, optic nerve, oral cavity, parotid gland, pharyngeal constrictor muscle (PCM), spinal cord, submandibular gland, temporal lobe, and temporomandibular joint (TMJ). All contouring tasks were performed using the Eclipse treatment planning system (version 15.6, Varian Medical Systems, Palo Alto, CA, USA). The full suite of tools available in the “Contouring” application, including interpolation, flood fill, and Boolean operators, was permitted to facilitate accurate and efficient segmentation.
DLC-u/DLC-m
A workstation equipped with a prototype of the DLC system was supplied by GE Healthcare for research purposes. This system employs a 2D U-Net convolutional neural network specifically configured for DLC applications. Such approaches have previously been utilized in magnetic resonance imaging-based DLC, as documented by Ruskó et al. (24). Throughout this investigation, the DLC system was tasked with generating the following anatomical structures: brainstem, chiasma, cochlea, eye, lacrimal gland, lens, mandible, optic nerve, parotid gland, PCM, pituitary gland, submandibular gland, temporal lobe, and thyroid gland. Although additional structures such as the spinal cord, esophagus, and oral cavity are available in the newer versions of the DLC system, they were not included in this study. Of the generated structures, the cochlea and temporal lobe were excluded from subsequent analysis due to incomplete data archiving. Three additional structures generated by the DLC system (lacrimal gland, pituitary gland, and thyroid gland) had no corresponding manual contours and were therefore not included in the comparison. The remaining nine structure types common to both the DLC and MC sets, brainstem, optic chiasm, eye, lens, mandible, optic nerve, parotid gland, PCM, and submandibular gland, constituted the evaluable cohort. Because five of these structures are bilateral organs, a total of 14 individual OARs were analyzed. After the automatic generation of the unmodified DLC (DLC-u) structure sets on the workstation, these sets were transferred to the Eclipse planning system to create modified structure sets (DLC-m). The modifications to achieve a clinically usable level were performed by the same radiation oncologist responsible for MC. One important consideration is that modifications are not intended to correct every error in the DLC-u structure sets. Instead, clinical decisions regarding whether to adjust a contour take into account multiple factors, including the distance of the OARs from high-dose regions and the potential severity of radiation therapy-related complications.
Objective evaluation
To compare three distinct OAR contouring strategies, the DSC and HD were employed as metrics. The DSC is a voxel-wise measure that quantifies the intersecting volume between two objects, defined as DSC = (2 |A ∩ B|) / (|A| + |B|). Conversely, the HD quantifies the maximum distance from any point in one set to the closest point in the other set and is expressed as HD = max{supb∈Binfa∈A d(b, a), supa∈Ainfb∈B d(b, a)}, where sup and inf denote the supremum and infimum, respectively, of the distances (25,26).
For the analysis of contouring time and to accurately reflect real-world time consumption, the OARs were categorized into three groups: Group A (structures above the skull base included in the DLC model), Group B (structures below the skull base included in the DLC model), and Group C (structures not included in the DLC model). A detailed list of these structures is provided in Table 2. The time required for MC was recorded for each group. Additionally, the time necessary for modifications was documented for Groups A and B, respectively. AI inference time was not systematically recorded and is not included in the reported contouring times.
Table 2
| Group A | Group B | Group C |
|---|---|---|
| Structures above skull base and included in the DLC model | Structures below skull base and included in the DLC model | Structures not included in the DLC model |
| Cochlea† | Brainstem | Brachial plexus |
| Eye | Mandible | Larynx |
| Lens | Parotid glands | Oral cavity |
| Optic chiasm | PCM | Spinal cord |
| Optic nerve | Submandibular gland | TMJ |
| Temporal lobe† |
†, excluded from DLC analysis due to incomplete data archiving. DLC, deep learning contours; PCM, pharyngeal constrictor muscle; TMJ, temporomandibular joint.
Subjective evaluation
For the DLC-u structure sets, an additional subjective 5-point scoring system, developed by clinical consensus among the participating radiation oncologists, was implemented to assess clinical usability. The operational definitions for this scoring system are as follows: 5 = Excellent (can be utilized directly without any deviations observed); 4 = Good (can be utilized directly with deviations that are clinically acceptable); 3 = Average (requires subtle modifications for clinical use, with an expected duration of less than 1 minute); 2 = Poor (requires minor modifications, anticipated to take less than 5 minutes); 1 = Very poor (requires major modifications, anticipated to exceed 5 minutes). Each DLC-u structure set was independently scored by the radiation oncologist responsible for that patient. This scoring system has not been independently validated in the literature.
Statistical analysis
All statistical analyses were conducted using R Statistical Software, versions 3.6.3 and 4.2.2 (27). The DSC and HD were computed for all three structure sets using the R programming environment and the pgirmess package (28). The correlation between various parameters was assessed using the Kendall rank correlation coefficient. To account for multiple comparisons, the Benjamini-Hochberg false discovery rate (FDR) correction was applied to all correlation analyses; both unadjusted and FDR-adjusted P values are reported. The mean times required for contouring and modification were analyzed using paired samples t-tests. Statistical significance was established at a P value of less than 0.05. Given the exploratory nature of this single-center validation study, a formal sample size calculation was not performed a priori. The sample of 50 consecutive patients was determined by the available cohort during the study period.
Ethical considerations
This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Institutional Review Board of the Chang Gung Medical Foundation (No. 201901223B0C503) and individual consent for this retrospective analysis was waived.
Results
Performance of DLC
Detailed results for the DSC and HD, comparing the DLC-u with MC, are presented in Table 3. For DSC values (Figure 1A), expressed as mean ± standard deviation, the results were favorable for the eyes [0.91±0.02 (right), 0.90±0.03 (left)], brainstem (0.87±0.03), parotid glands [0.84±0.07 (right), 0.84±0.04 (left)] and submandibular glands [0.80±0.15 (right), 0.79±0.16 (left)], with values either exceeding or approaching 0.8. Conversely, the optic chiasm (0.44±0.19), PCM (0.42±0.11), and mandible (0.41±0.21) showed suboptimal performance with DSC values below 0.5. Regarding HD values used as an evaluative parameter (Figure 1B), the lens [2.25±0.76 mm (right), 2.65±0.87 mm (left)] and eyes [3.00±0.83 mm (right), 2.93±0.72 mm (left)] demonstrated good outcomes with mean HD values less than 5 mm. In contrast, the parotid glands [13.43±5.80 mm (right), 11.68±4.84 mm (left)], PCM (20.11±6.62 mm), and mandible (27.49±8.56 mm) recorded higher distances. Subjective scoring results (Figure 1C) indicated that most organs reached a clinically usable level with only subtle modifications required (score ≥3), with the exceptions of the optic chiasm (2.44±0.86), optic nerves [2.92±0.49 (left), 2.86±0.61 (right)], and PCM (2.98±1.02), which scored lower.
Table 3
| Organ at risk | MC vs. DLC-u | DLC-u vs. DLC-m | MC vs. DLC-m | Subjective scores | |||||
|---|---|---|---|---|---|---|---|---|---|
| DSC | HD | DSC | HD | DSC | HD | ||||
| Brainstem | 0.87±0.03 | 6.37±1.55 | 0.97±0.02 | 4.71±2.16 | 0.87±0.02 | 6.22±1.72 | 3.46±0.81 | ||
| Optic chiasma | 0.44±0.19 | 5.19±1.90 | 0.55±0.26 | 4.54±2.21 | 0.52±0.18 | 4.68±2.07 | 2.44±0.86 | ||
| Left eye | 0.90±0.03 | 2.93±0.72 | 0.99±0.05 | 0.70±1.67 | 0.89±0.06 | 2.98±1.38 | 4.58±0.7 | ||
| Right eye | 0.91±0.02 | 3.00±0.83 | 0.99±0.04 | 0.71±1.63 | 0.90±0.04 | 3.16±1.37 | 4.52±0.65 | ||
| Left lens | 0.74±0.07 | 2.65±0.87 | 0.98±0.09 | 0.45±1.20 | 0.73±0.10 | 2.70±1.03 | 4.49±0.79 | ||
| Right lens | 0.76±0.06 | 2.25±0.76 | 0.99±0.08 | 0.36±1.17 | 0.75±0.09 | 2.21±0.80 | 4.61±0.7 | ||
| Mandible | 0.41±0.21 | 27.49±8.56 | 0.98±0.02 | 10.16±10.89 | 0.42±0.21 | 28.59±11.24 | 3.58±0.78 | ||
| Left optic nerve | 0.64±0.13 | 6.40±4.30 | 0.82±0.16 | 7.91±6.75 | 0.68±0.11 | 6.17±6.67 | 2.92±0.49 | ||
| Right optic nerve | 0.64±0.09 | 8.51±6.69 | 0.83±0.11 | 10.83±8.01 | 0.68±0.07 | 5.08±4.40 | 2.86±0.61 | ||
| Left parotid gland | 0.84±0.04 | 11.68±4.84 | 0.99±0.02 | 4.21±4.59 | 0.85±0.04 | 10.81±4.67 | 3.82±0.77 | ||
| Right parotid gland | 0.84±0.07 | 13.43±5.80 | 0.98±0.05 | 3.83±4.7 | 0.84±0.05 | 12.38±5.26 | 3.84±0.91 | ||
| Pharyngeal constrictor muscle | 0.42±0.11 | 20.11±6.62 | 0.80±0.22 | 13.74±40.19 | 0.47±0.08 | 23.64±38.84 | 2.98±1.02 | ||
| Left submandibular gland | 0.79±0.16 | 7.51±3.43 | 0.95±0.13 | 3.93±4.73 | 0.81±0.14 | 7.57±3.61 | 3.62±1.25 | ||
| Right submandibular gland | 0.80±0.15 | 7.89±3.79 | 0.93±0.15 | 4.11±4.53 | 0.81±0.11 | 7.6±3.75 | 3.62±1.19 | ||
Data are presented as mean ± standard deviation. DLC-m, deep learning contouring with modification; DLC-u, unmodified deep learning contouring; DSC, Dice similarity coefficient; HD, Hausdorff distance; MC, manual contouring.
To simulate real-world clinical applications, the DLC-u structure sets were adjusted to create the DLC-m structure sets. The extent of these modifications is also quantified by employing both the DSC and HD, as detailed in Table 3. With respect to DSC, the most notable changes were observed for the optic chiasm (0.55±0.26), PCM (0.80±0.22), and optic nerves [0.82±0.16 (left), 0.83±0.11 (right)]. Regarding HD, the largest discrepancies were noted for PCM (13.74±40.19 mm), optic nerves [7.91±6.75 mm (left), 10.83±8.01 mm (right)], and mandible (10.16±10.89 mm). A similar comparison between the clinically approved structure sets, DLC-m and MC, is detailed in Table 3.
Correlation between objective and subjective parameters
To further illustrate the discrepancies between subjective assessments and objective measurements, scatter plots depicting correlations of subjective scores with DSC and HD values are provided in Figures 2,3, respectively. Additionally, Kendall’s tau correlation analyses were conducted (Table 4). After Benjamini-Hochberg FDR correction, statistically significant correlations between the DSC values and the subjective scores were retained for the right parotid gland (τ=0.39, q=0.014) and optic chiasma (τ=0.34, q=0.035). Several other structures showed nominally significant unadjusted P values (mandible, PCM, right submandibular gland; all P<0.05) but did not survive FDR correction. Regarding the correlation between HD and subjective scores, significant correlations after FDR correction were observed for the right optic nerve (τ=−0.42, q<0.014) and left parotid gland (τ=−0.32, q=0.028). Additional structures with nominally significant unadjusted P values (left lens, left submandibular gland, optic chiasma) did not reach significance after correction.
Table 4
| Organ at risk | DSC | HD | |||||
|---|---|---|---|---|---|---|---|
| τ | P value | q value | τ | P value | q value | ||
| Brainstem | 0.10 | 0.39 | 0.493 | −0.21 | 0.06 | 0.147 | |
| Optic chiasma | 0.34 | 0.005 | 0.035* | −0.25 | 0.042 | 0.118 | |
| Left eye | −0.09 | 0.46 | 0.533 | −0.05 | 0.68 | 0.731 | |
| Right eye | 0.03 | 0.83 | 0.869 | −0.03 | 0.80 | 0.798 | |
| Left lens | −0.02 | 0.87 | 0.869 | −0.29 | 0.02 | 0.07 | |
| Right lens | 0.19 | 0.10 | 0.179 | −0.12 | 0.34 | 0.43 | |
| Mandible | 0.24 | 0.03 | 0.112 | −0.11 | 0.31 | 0.43 | |
| Left optic nerve | 0.18 | 0.14 | 0.212 | −0.15 | 0.20 | 0.345 | |
| Right optic nerve | 0.13 | 0.25 | 0.353 | −0.42 | <0.001 | <0.014* | |
| Left parotid gland | 0.21 | 0.07 | 0.138 | −0.32 | 0.004 | 0.028* | |
| Right parotid gland | 0.39 | 0.001 | 0.014* | −0.19 | 0.08 | 0.168 | |
| Pharyngeal constrictor muscle | 0.24 | 0.03 | 0.112 | 0.06 | 0.58 | 0.677 | |
| Left submandibular gland | 0.24 | 0.053 | 0.124 | −0.28 | 0.02 | 0.07 | |
| Right submandibular gland | 0.24 | 0.043 | 0.12 | −0.12 | 0.33 | 0.43 | |
*, q<0.05 after Benjamini-Hochberg false discovery rate correction. DSC, Dice similarity coefficient; HD, Hausdorff distance.
Contouring time
In the MC of all OARs, including those typically involved in nasopharyngeal cancer (NPC) treatment scenarios (Group A + B + C), the average contouring duration was recorded as 35.3±7.2 minutes. This duration was significantly reduced to 24.7±5.2 minutes when utilizing the DLC-m approach, yielding a mean reduction of 10.6 minutes (P<0.001, Figure 4). Excluding structures above the skull base, as is common in most HNC treatments (Group B + C), a similar reduction in time was observed, from 25.7±5.0 to 20.6±5.2 minutes (mean difference 5.1 minutes, P<0.001). When omitting structures not included in the DLC model from the analysis (Group A + B), the contouring time decreased significantly, from 21.8±4.7 to 11.2±2.9 minutes (mean difference 10.6 minutes, P<0.001).
Discussion
In this study, we evaluated the performance of a prototype DLC system within the context of HNC RT. Per-organ DSC ranged from 0.41 to 0.91 and subjective scoring indicated that, across all structure-level evaluations, a majority (86.3%) received a score ≥3, corresponding to clinically usable contours after no or only subtle modifications. For the structures available in the DLC model, the human-in-the-loop workflow achieved a 48.6% reduction in contouring time (21.8 → 11.2 minutes), corresponding to an overall 30.0% reduction (35.3 → 24.7 minutes) when all OARs required for treatment planning, including those not covered by the DLC model, were considered. Although the prototype included only a limited number of organs at the time of analysis, it demonstrated generally favorable performance in both subjective and objective assessments. While this deep learning-based system does not yet replace expert manual delineation in HNC RT, it has reached a level of clinical usability that can substantially assist medical personnel. A major strength of our study lies in the diversity of clinical scenarios represented, which supports the generalizability of DLC in real-world clinical settings. More than half of the patients included had stage III or IV disease, and 14% had undergone surgery prior to RT. The anatomical distortions resulting from either surgical intervention or locally advanced disease pose significant challenges in the contouring of OARs (14). Moreover, our cohort included a higher proportion of patients with NPC, which requires larger RT fields and greater consideration of intracranial structures compared to oropharyngeal and laryngeal cancers. This composition further enhances the applicability of our findings to complex treatment settings that are underrepresented in Western-centric datasets.
Overall, the contour quality observed in our study is consistent with findings reported in several previous studies (29-33). The system exhibited suboptimal performance in delineating the optic chiasm, PCM, and mandible, as reflected by relatively low DSC values. Precisely delineating the boundaries of neural structures, such as the optic apparatus, brainstem, and spinal cord, remains challenging even for experienced radiation oncologists due to the relatively low imaging contrast between neural tissue and surrounding cerebrospinal fluid, as well as imaging artifacts introduced by adjacent bony structures. In their study utilizing a commercial 3D U-Net system, Ng et al. reported DSC values ranging from 0.31 to 0.88 and specifically noted suboptimal contouring accuracy for structures including the cochleae, larynx, optic chiasm, oral cavity, and spinal cord (32). Similarly, Zhong et al. documented suboptimal DSC values for the optic chiasm (0.46) and optic nerves (0.44 and 0.51 for the right and left sides, respectively) (31). Regarding the suboptimal performance of PCM, Koo et al. (30), employing a fully convolutional neural network that integrates U-Net and V-Net architectures, reported an overall mean DSC of 0.81 (range, 0.66–0.89), with the lowest DSC observed for the PCM at 0.66. Since the PCM lacks clearly defined boundaries on CT imaging and is instead delineated as an arbitrary 3-mm thick margin surrounding the aerodigestive tract, considerable variability in PCM contouring is frequently encountered in both clinical and research settings (34,35). Consequently, a heightened level of uncertainty has led to increased tolerance among clinical personnel for inaccuracies in contour delineation. On the other hand, the suboptimal performance observed for the mandible represents a relatively unique finding (16). Reviewers noted a common omission of the upper mandible region near the condyle. It is hypothesized that because the training dataset was derived from clinical treatments, parts distant from high-dose regions might have been omitted to expedite the contouring process. Although we do not have direct access to the original training dataset to confirm this hypothesis, this finding underscores the importance of high-quality training datasets and highlights a common dilemma faced in machine learning-based applications deployed in clinical scenarios. Furthermore, dental artifacts, despite the application of a metal artifact reduction algorithm, may have degraded segmentation accuracy for the mandible and adjacent structures. The inclusion of postoperative patients (14%), whose surgically altered anatomy deviates from the training data, likely also contributed to the large standard deviations observed. We did not perform a formal subgroup analysis of failure modes (e.g., DSC stratified by presence of dental artifacts or prior surgery) owing to limited subgroup sizes; the mechanisms proposed above should therefore be regarded as observational hypotheses rather than statistically demonstrated effects.
The weak to moderate correlation between objective metrics and subjective evaluations is noteworthy. This discrepancy is partly attributable to the inherent volume dependence of the DSC: for small structures such as the optic chiasm, even a minor boundary deviation can produce a disproportionately low DSC, whereas the same absolute deviation in a large organ like the parotid gland has minimal impact on the metric. In some cases with suboptimal objective measurements, a considerable proportion of the structures were still deemed clinically acceptable upon subjective review, such as PCM and mandible, which are generally more tolerant of contouring variability. Conversely, certain structures required further modification for clinical use despite exhibiting favorable objective metrics. This is particularly true for neural structures such as the brainstem, which require highly precise delineation due to their radiosensitivity and the potential for unacceptable radiation therapy-related toxicities. Similarly, the objective comparison between MC and DLC-m offers a similar perspective. Although both structure sets were deemed clinically acceptable, the DSCs were still notably below one, and the HDs remained substantial. This observation also supports that minor geometric discrepancies may be deemed acceptable depending on the organ involved and its proximity to high-dose regions.
Theoretically, deeper or more complex architectures, such as 3D-based deep learning models (16), could be employed to more accurately approximate the ground truth. However, in practice, the substantial computational overhead associated with these models presents a significant barrier. To address this, the integration of models with varying levels of complexity should be considered to achieve clinically acceptable results more efficiently. Lightweight models may be appropriate for less clinically critical OARs, where minor inaccuracies are tolerable. Conversely, more complex models may be reserved for critical structures that demand high geometric precision. Such a hybrid approach offers a pragmatic solution for balancing computational efficiency with contouring accuracy, thereby enhancing the feasibility of DLC implementation in time-sensitive clinical workflows, such as daily adaptive RT.
There are several limitations in this study. First, although this study was conducted at a tertiary medical center (Linkou Chang Gung Memorial Hospital, Taoyuan, Taiwan), its single-center design and relatively small cohort may limit the generalizability of the findings. Furthermore, as the DLC model was developed and trained by the vendor on an external dataset, the training data may not fully represent the anatomical variability and contouring conventions of our institution or other centers, potentially limiting the generalizability of the observed performance. Also, the DLC model did not encompass all OARs required for HNC RT planning, thereby limiting the scope of the evaluation. Second, several design choices constrain the interpretability of the reference standard. Each patient’s manual contours were generated by a single one of the three radiation oncologists, without a formal consensus mechanism (e.g., STAPLE) or quantitative assessment of inter- or intra-observer variability; consequently, the magnitude of observer noise in the reference standard is unknown, and some proportion of the DLC-versus-MC discrepancies we report could reflect individual contouring preference rather than true algorithmic error. In addition, the same clinician who performed MC also post-processed the DLC outputs, which may introduce anchoring bias: clinicians are known to be cognitively predisposed to accept pre-generated contours that appear “good enough” rather than to rigorously correct them. A blinded, independent modifier was not feasible within the scope of this study, and we therefore cannot exclude that the DLC-m/MC concordance is partially driven by within-observer consistency. Another important limitation concerns the subjective evaluation instrument. The 5-point scoring scale used in this study was developed internally by clinical consensus and has not been independently validated in prior literature. In addition, each DLC-u structure set was scored only by the radiation oncologist responsible for that patient, so inter-rater reliability (e.g., Cohen’s kappa) could not be computed. Accordingly, all subjective-score–derived statements should be regarded as exploratory rather than confirmatory. Third, and most importantly, we did not perform a dosimetric evaluation. Geometric agreement (DSC, HD) does not linearly translate into dosimetric safety: a small boundary error in a high-dose-gradient region (e.g., brainstem, optic chiasm) may be clinically critical, whereas a larger error in a low-dose region may be negligible. Without comparing dose-volume histograms derived from MC-based and DLC-based contours on the same treatment plans, the clinical safety of deploying the DLC-u or DLC-m contours cannot be established. All “clinical utility” statements in this work should therefore be read as being restricted to geometric similarity and time efficiency, not dosimetric equivalence. Fourth, although our cohort spanned diverse HNC subsites and both definitive and postoperative settings, subgroup sizes were too small to support stratified inference (e.g., postoperative vs. definitive, presence vs. absence of dental artifacts). No a priori power calculation was performed; our sample of 50 consecutive patients reflected the available cohort during the study period, and while it is comparable in size to prior single-institution DLC validation studies, we acknowledge that such a sample is underpowered to detect clinically meaningful differences for individual OARs with high anatomical variability. Per-OAR estimates in this study should be regarded as hypothesis-generating and interpreted together with their standard deviations rather than as point estimates of true performance.
Conclusions
In this exploratory single-center validation, the evaluated DLC prototype reduced contouring time by 48.6% for the structures covered by the model, corresponding to an overall 30.0% reduction in total OAR contouring time. A total of 86% of individual structure-level subjective evaluations reached a score ≥3, indicating clinical usability with no or only minor edits. However, performance was suboptimal for several small or low-contrast structures, including the optic chiasm (DSC = 0.44), PCM (DSC = 0.42), and mandible (DSC = 0.41), for which substantial manual intervention remained necessary. Moreover, the weak concordance between geometric metrics and subjective clinical judgement suggests that DSC and HD alone are insufficient to fully capture clinical acceptability. Because dosimetric impact was not assessed in this study, claims regarding clinical safety remain preliminary. The system should be considered a supportive tool within a human-in-the-loop workflow that requires careful expert review, rather than a validated substitute for manual delineation.
Acknowledgments
None.
Footnote
Data Sharing Statement: Available at https://tro.amegroups.com/article/view/10.21037/tro-25-44/dss
Peer Review File: Available at https://tro.amegroups.com/article/view/10.21037/tro-25-44/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tro.amegroups.com/article/view/10.21037/tro-25-44/coif). C.W.L. and I.L.S. are employees of GE Healthcare, the vendor of the prototype deep-learning contouring system evaluated in this study. They participated in data collection and provided technical support for the research workstation but were not involved in the independent statistical analysis or the drafting of the study’s conclusions. The other authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Institutional Review Board of the Chang Gung Medical Foundation (No. 201901223B0C503) and individual consent for this retrospective analysis was waived.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Sung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin 2021;71:209-49. [Crossref] [PubMed]
- Li B, Li D, Lau DH, et al. Clinical-dosimetric analysis of measures of dysphagia including gastrostomy-tube dependence among head and neck cancer patients treated definitively by intensity-modulated radiotherapy with concurrent chemotherapy. Radiat Oncol 2009;4:52. [Crossref] [PubMed]
- Schwartz DL, Hutcheson K, Barringer D, et al. Candidate dosimetric predictors of long-term swallowing dysfunction after oropharyngeal intensity-modulated radiotherapy. Int J Radiat Oncol Biol Phys 2010;78:1356-65. [Crossref] [PubMed]
- Beetz I, Schilstra C, van der Schaaf A, et al. NTCP models for patient-rated xerostomia and sticky saliva after treatment with intensity modulated radiotherapy for head and neck cancer: the role of dosimetric and clinical factors. Radiother Oncol 2012;105:101-6. [Crossref] [PubMed]
- Gomez DR, Estilo CL, Wolden SL, et al. Correlation of osteoradionecrosis and dental events with dosimetric parameters in intensity-modulated radiation therapy for head-and-neck cancer. Int J Radiat Oncol Biol Phys 2011;81:e207-13. [Crossref] [PubMed]
- Fatigue following radiation therapy in nasopharyngeal cancer survivors: A dosimetric analysis incorporating patient report and observer rating. Radiother Oncol 2019;133:35-42.
- Vorwerk H, Zink K, Schiller R, et al. Protection of quality and innovation in radiation oncology: the prospective multicenter trial the German Society of Radiation Oncology (DEGRO-QUIRO study). Evaluation of time, attendance of medical staff, and resources during radiotherapy with IMRT. Strahlenther Onkol 2014;190:433-43.
- Harari PM, Song S, Tomé WA. Emphasizing conformal avoidance versus target definition for IMRT planning in head-and-neck cancer. Int J Radiat Oncol Biol Phys 2010;77:950-8. [Crossref] [PubMed]
- Apolle R, Appold S, Bijl HP, et al. Inter-observer variability in target delineation increases during adaptive treatment of head-and-neck and lung cancer. Acta Oncol 2019;58:1378-85. [Crossref] [PubMed]
- Nelms BE, Tomé WA, Robinson G, et al. Variations in the contouring of organs at risk: test case from a patient with oropharyngeal cancer. Int J Radiat Oncol Biol Phys 2012;82:368-78. [Crossref] [PubMed]
- van der Veen J, Gulyban A, Willems S, et al. Interobserver variability in organ at risk delineation in head and neck cancer. Radiat Oncol 2021;16:120. [Crossref] [PubMed]
- Han X, Hoogeman MS, Levendag PC, et al. Atlas-based auto-segmentation of head and neck CT images. Med Image Comput Comput Assist Interv 2008;11:434-41. [Crossref] [PubMed]
- Teguh DN, Levendag PC, Voet PW, et al. Clinical validation of atlas-based auto-segmentation of multiple target volumes and normal tissue (swallowing/mastication) structures in the head and neck. Int J Radiat Oncol Biol Phys 2011;81:950-7. [Crossref] [PubMed]
- Daisne JF, Blumhofer A. Atlas-based automatic segmentation of head and neck organs at risk and nodal target volumes: a clinical validation. Radiat Oncol 2013;8:154. [Crossref] [PubMed]
- Ibragimov B, Xing L. Segmentation of organs-at-risks in head and neck CT images using convolutional neural networks. Med Phys 2017;44:547-57. [Crossref] [PubMed]
- Liu P, Sun Y, Zhao X, et al. Deep learning algorithm performance in contouring head and neck organs at risk: a systematic review and single-arm meta-analysis. Biomed Eng Online 2023;22:104. [Crossref] [PubMed]
- Ronneberger O, Fischer P, Brox T. U-Net: Convolutional networks for biomedical image segmentation. In: Navab N, Hornegger J, Wells W, Frangi A, editors. Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. Lecture Notes in Computer Science, vol 9351. Springer; 2015:234-41.
- Lustberg T, van Soest J, Gooding M, et al. Clinical evaluation of atlas and deep learning based automatic contouring for lung cancer. Radiother Oncol 2018;126:312-7. [Crossref] [PubMed]
- Ahn SH, Yeo AU, Kim KH, et al. Comparative clinical evaluation of atlas and deep-learning-based auto-segmentation of organ structures in liver cancer. Radiat Oncol 2019;14:213. [Crossref] [PubMed]
- Kim H, Jung J, Kim J, et al. Abdominal multi-organ auto-segmentation using 3D-patch-based deep convolutional neural network. Sci Rep 2020;10:6204. [Crossref] [PubMed]
- Cha E, Elguindi S, Onochie I, et al. Clinical implementation of deep learning contour autosegmentation for prostate radiotherapy. Radiother Oncol 2021;159:1-7. [Crossref] [PubMed]
- Zabel WJ, Conway JL, Gladwish A, et al. Clinical Evaluation of Deep Learning and Atlas-Based Auto-Contouring of Bladder and Rectum for Prostate Radiation Therapy. Pract Radiat Oncol 2021;11:e80-9. [Crossref] [PubMed]
- Yom SS, Torres-Saavedra P, Caudell JJ, et al. Reduced-Dose Radiation Therapy for HPV-Associated Oropharyngeal Carcinoma (NRG Oncology HN002). J Clin Oncol 2021;39:956-65. [Crossref] [PubMed]
- Ruskó L, Capala ME, Czipczer V, et al. Deep-Learning-based Segmentation of Organs-at-Risk in the Head for MR-assisted Radiation Therapy Planning. BIOIMAGING 2021;2:31-43.
- Yang J, Sharp GC, Gooding MJ, editors. Auto-segmentation for radiation oncology: state of the art. CRC Press; 2021.
- Rote G. Computing the minimum Hausdorff distance between two point sets on a line under translation. Inf Process Lett 1991;38:123-7.
- R Core Team. R: A Language and Environment for Statistical Computing. 3.6.3 ed. Vienna, Austria: R Foundation for Statistical Computing; 2020.
- Giraudoux P. pgirmess: Spatial analysis and data mining for field ecologists. R package version 1.7.1. 2021.
- van Dijk LV, Van den Bosch L, Aljabar P, et al. Improving automatic delineation for head and neck organs at risk by Deep Learning Contouring. Radiother Oncol 2020;142:115-23. [Crossref] [PubMed]
- Koo J, Caudell JJ, Latifi K, et al. Comparative evaluation of a prototype deep learning algorithm for autosegmentation of normal tissues in head and neck radiotherapy. Radiother Oncol 2022;174:52-8. [Crossref] [PubMed]
- Zhong Y, Yang Y, Fang Y, et al. A Preliminary Experience of Implementing Deep-Learning Based Auto-Segmentation in Head and Neck Cancer: A Study on Real-World Clinical Cases. Front Oncol 2021;11:638197. [Crossref] [PubMed]
- Ng CKC, Leung VWS, Hung RHM. Clinical Evaluation of Deep Learning and Atlas-Based Auto-Contouring for Head and Neck Radiation Therapy. App Sci 2022;12:11681.
- Lucido JJ, DeWees TA, Leavitt TR, et al. Validation of clinical acceptability of deep-learning-based automated segmentation of organs-at-risk for head-and-neck radiotherapy treatment planning. Front Oncol 2023;13:1137803. [Crossref] [PubMed]
- Petkar I, McQuaid D, Dunlop A, et al. Inter-Observer Variation in Delineating the Pharyngeal Constrictor Muscle as Organ at Risk in Radiotherapy for Head and Neck Cancer. Front Oncol 2021;11:644767. [Crossref] [PubMed]
- Brouwer CL, Steenbakkers RJ, Bourhis J, et al. CT-based delineation of organs at risk in the head and neck region: DAHANCA, EORTC, GORTEC, HKNPCSG, NCIC CTG, NCRI, NRG Oncology and TROG consensus guidelines. Radiother Oncol 2015;117:83-90. [Crossref] [PubMed]
Cite this article as: Chen PJ, Lee CH, Hung SP, Lee CW, Shih IL, Cheng AJ, Chang JTC. Geometric accuracy and time efficiency evaluation of a deep learning auto-segmentation system for head and neck radiotherapy: a single-center validation study. Ther Radiol Oncol 2026;10:8.

