SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

CVPR 2026 Ye-Chan Kim, SeungJu Cha, Si-Woo Kim, MinJu Jeon, HyunGee Kim, Dong-Jin Kim

Abstract

SAIL tackles weakly-supervised dense video captioning through similarity-aware guidance and inter-caption augmentation. The framework reduces reliance on dense temporal annotations while maintaining strong captioning quality.