A benchmark and taxonomy for phonemizing noisy user-generated text, paired with a compositional approach for irregular spellings, abbreviations, and code-mixing.
Publications
Research on multimodal learning, language models, and data-centric AI.
Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
A retrieval-augmented dense video captioning approach that injects supervised saliency signals into both the retrieval and generation stages.
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
Tackles weakly-supervised dense video captioning through similarity-aware guidance and inter-caption augmentation.
Cap4Bridge: Caption-Guided Cross-Modal Contextualization with Stochastic Augmentation for Text-Video Retrieval
Bridges the text-video modality gap by using generated captions as cross-modal context, enriched through stochastic augmentation during training.
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
A dense video captioning framework that reweights video frames by saliency and adaptively retrieves relevant captions at inference time.
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
Refines noisy synthetic image-caption datasets through a one-to-many mapping that re-aligns each image with its best-matching captions.