Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning

CVPR 2026 Seunghee Choi, MinJu Jeon, Hyunwoo Oh, Jihwan Lee, Dong-Jin Kim

Abstract

A retrieval-augmented dense video captioning approach that injects supervised saliency signals into both the retrieval and generation stages. The saliency guidance helps the model focus on temporally important moments, producing more grounded and informative captions.