Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning

EMNLP 2025 (Long, Main) MinJu Jeon, Si-Woo Kim, Ye-Chan Kim, HyunGee Kim, Dong-Jin Kim

Abstract

A dense video captioning framework that reweights video frames by saliency and adaptively retrieves relevant captions at inference time. By concentrating supervision on semantically important moments and grounding generation in retrieved context, Sali4Vid produces more accurate and temporally localized descriptions.