115th Annual Meeting of the United-States-and-Canadian-Academy-of-Pathology (USCAP), San-Antonio, Kuzey Mariana Adaları, 21 - 26 Mart 2026, cilt.106, (Özet Bildiri)
PD-L1 is a predictive biomarker in urothelial carcinoma. However, interobserver variability in PD-L1 scoring is a major challenge, particularly among less-experienced pathologists. Digital pathology algorithms have the potential to reduce variability and align junior assessments with expert standards.
Design
A total of 21 urothelial carcinoma cases with both PD-L1 IHC were included. Two senior pathologists identified hot-spot regions and measured tumor proportion score (TPS%) for PD-L1 Three pathology residents (R1, R2, R3) first evaluated each slide independently without assistance, and with assistance of the algorithms’ masks and scores. Accuracy was quantified using mean absolute error (MAE), root mean square error (RMSE), intraclass correlation coefficient (ICC), Cohen’s kappa (using clinical cut-offs: <1%, 1–49%, ≥50%), and Bland–Altman analysis. Paired statistical tests assessed differences between algorithm-assisted and unassisted resident performance.
Results
The results of the algorithm (Figure) were similar to those of the senior pathologists (T-test = 0.16). For R1, MAE was 2.71 without the algorithm and 2.43 with it. RMSE decreased from 5.24 to 4.56; ICC increased from 0.986 to 0.990; and the Kappa decreased from 0.92 to 0.84. For R2, the MAE improved from 3.76 to 2.86, the RMSE from 7.42 to 6.18 and the ICC from 0.974 to 0.983; meanwhile, the Kappa was similar as 0.76. For R3, the MAE was 2.19 without the algorithm and 2.52 with it, the RMSE was 4.71 versus 4.90, the ICC was 0.989 versus 0.989, and the Kappa was 0.84. Paired t-tests comparing resident and senior pathologist scores yielded p-values >0.05 across all residents.
Figure 1 - 1365
Conclusions
Algorithm assistance reduced error and increased correlation for R1 and R2. However, R3 showed no additional benefit due to the high level of initial agreement. These results suggest that algorithm support could have the greatest impact on residents with lower baseline accuracy, improving reproducibility in PD-L1 scoring. However, larger datasets are needed to confirm statistical significance and clinical utility.