IEEE Robotics and Automation Letters Accepted July 2026
Diff2DGS Reliable Reconstruction of Occluded Surgical Scenes via 2D Gaussian Splatting
Temporally guided instrument removal and deformable 2D Gaussian Splatting for appearance-faithful, geometry-aware reconstruction of dynamic surgical scenes.
Department of Computer Science and UCL Hawkes Institute, University College London
- EndoNeRF PSNR
- 38.02 dB
- StereoMIS PSNR
- 34.40 dB
- Rendering
- 200+ FPS
- Evaluation
- 3 datasets
Why Diff2DGS
Reconstruction that remains reliable beyond the input view
Image-space quality alone can hide severe geometric artifacts. Diff2DGS restores instrument-occluded tissue before reconstruction and explicitly balances appearance with depth supervision.
TL;DR
Diff2DGS restores tissue hidden by surgical instruments using temporally guided diffusion, then reconstructs the dynamic scene with deformable 2D Gaussian Splatting and adaptive depth supervision.
Method
A two-stage route from occluded video to dynamic 3D tissue
Temporal video priors guide diffusion-based instrument removal and tissue inpainting across frames.
Results
Inspect appearance, occlusion recovery, and geometry
Use the view selector to move through the main evidence reported across EndoNeRF, StereoMIS, and SCARED.
Qualitative reconstruction
Diff2DGS recovers cleaner tissue appearance in regions hidden by surgical instruments across StereoMIS and EndoNeRF.
Quantitative summary
Real-time rendering with stronger image fidelity
Depth RMSE on EndoNeRF and StereoMIS is measured against stereo-depth references. Rendering speed is measured on a single NVIDIA Tesla V100.
| Dataset | PSNR | SSIM | RMSE | FPS |
|---|---|---|---|---|
| EndoNeRF | 38.02 dB | 96.44% | 2.51 | 232.29 |
| StereoMIS | 34.40 dB | 90.72% | 3.79 | 212.32 |
Video demo
Occlusion removal before dynamic reconstruction
The surgical input, Deform3DGS reconstruction, and Diff2DGS reconstruction are shown on a shared timeline for direct comparison.
Paper
Abstract
Real-time reconstruction of deformable surgical scenes is vital for advancing robotic surgery, improving intraoperative guidance, and enabling automation. Recent methods achieve dense reconstructions from da Vinci robotic surgery videos, with Gaussian Splatting offering real-time performance via graphics acceleration. However, reconstruction quality in occluded regions remains limited, and depth accuracy has not been fully assessed, as benchmarks like EndoNeRF and StereoMIS lack 3D ground truth.
We propose Diff2DGS, a two-stage framework for reliable 3D reconstruction of occluded surgical scenes. First, a diffusion-based video module with temporal priors inpaints tissue occluded by instruments with high spatiotemporal consistency. Second, we adapt 2D Gaussian Splatting with a Learnable Deformation Model to capture dynamic tissue deformation and anatomical geometry, and introduce adaptive depth weight to improve geometric fidelity. We further extend evaluation beyond image-quality metrics by performing quantitative depth analysis on the SCARED dataset.
Diff2DGS outperforms state-of-the-art methods in appearance quality, reaching 38.02 dB PSNR on EndoNeRF and 34.40 dB on StereoMIS. Our experiments also show that optimizing image quality alone does not necessarily ensure accurate 3D geometry.
- Computer Vision for Medical Robotics
- Surgical Robotics: Laparoscopy
- Surgical Robotics: Planning
Reference
Cite Diff2DGS
The paper is accepted for publication in IEEE Robotics and Automation Letters. Volume, issue, pages, and DOI will be added after IEEE Xplore posting.
@article{song2026diff2dgs,
author = {Song, Tianyi and Stoyanov, Danail and
Mazomenos, Evangelos and Vasconcelos, Francisco},
title = {{Diff2DGS}: Reliable Reconstruction of Occluded
Surgical Scenes via 2D Gaussian Splatting},
journal = {IEEE Robotics and Automation Letters},
year = {2026},
note = {Accepted for publication}
}
Acknowledgments
Supported by EPSRC OASIS [UKRI145], DSIT and the Royal Academy of Engineering Chair in Emerging Technologies programme, EPSRC HuMIRoS [EP/Z534754/1], and the UCL Centre for Digital Innovation Scholarship.