Appendix A: Ethical approval report

The scanned copy of the formal ethical approval report for the endoscopic video dataset is shown in Fig. 8.

Fig. 8Fig. 8

The scanned copy of the formal ethical approval report for the endoscopic video dataset.

Appendix B: Additional experimental details

This appendix provides additional details on the clinical dataset, data splitting strategy, synthetic degradation settings, training configurations, downstream VSR training settings, and t-SNE visualization settings used in the experiments.

Synthetic dataset information

The information of the dataset used in “Evaluating simulated LR images” and “VSR on the synthetic dataset” sections is shown in Table 9.

Table 9 Statistics of the collected clinical endoscopic video dataset.Data splitting strategy

The data splitting strategy used in the experiments is summarized in Table 10. We divided the 200 selected HR video sequences into two non-overlapping subsets. The first subset was used as the source-domain HR data for degradation learning. The second subset was processed by different synthetic degradation settings to construct the target-domain LR data for adversarial degradation learning. In addition, 20 extra video sequences were reserved for testing. The split was performed at the video-sequence level to avoid information leakage.

Table 10 Data splitting strategy used in the experiments.Synthetic degradation settings

The detailed synthetic degradation settings are reported in Table 11. This synthetic benchmark is designed to cover several typical degradation factors that may appear in endoscopic videos, including interpolation degradation, isotropic Gaussian blur, direction-dependent anisotropic blur, motion blur, and noise. All degradation settings are conducted under the \(\times 4\) upscaling factor.

Table 11 Synthetic degradation settings used in the experiments.UDLM training settings

The training settings of the proposed UDLM are summarized in Table 12. During training, each input sample consists of a video clip with 10 consecutive frames. We crop HR patches of size \(256 \times 256\) from the original HR frames. Under the \(\times 4\) super-resolution setting, the corresponding LR patch size is \(64 \times 64\). The model is first trained with the low-frequency loss during the warm-up stage. After that, the adaptive data loss is introduced, and the adaptive kernel is periodically updated according to the current behavior of the degradation generator.

Table 12 Training settings of the proposed UDLM.Loss weights

The loss weights used for UDLM training are listed in Table 13. The low-frequency loss and adaptive data loss share the same data-loss weight. After the warm-up stage, the data constraint is switched from the low-frequency loss to the adaptive data loss.

Table 13 Loss weights used for UDLM training.Training settings of downstream VSR models

The training settings of the downstream VSR models are summarized in Table 14. For fair comparison, all pseudo-paired data generated by different degradation learning methods are used to train the downstream VSR models under the same configuration.

Table 14 Training settings of downstream VSR models.Real clinical dataset statistics

The statistics of the real clinical endoscopic dataset used in “Performance evaluation in real-world endoscopic scenarios” section are shown in Table 15.

Table 15 Statistics of the real clinical endoscopic dataset used in “Performance evaluation in real-world endoscopic scenarios” section.t-SNE visualization settings

The detailed settings of the t-SNE visualization experiment in real endoscopic scenarios are summarized in Table 16.

Table 16 Settings of the t-SNE visualization experiment.