Photo of MinJu Jeon

MinJu Jeon

LG AI Research · EXAONE Lab

AI Researcher · Multimodal & Language Models · Data-Centric Training

  • About
  • Publications
  • Projects
  • CV

#data-centric

Content tagged with "data-centric"

Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach
EMNLP 2026 (Findings, Long) MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee 2026-11-01
#Data-Centric #Multilingual Speech #Grapheme-to-Phoneme

A benchmark and taxonomy for phonemizing noisy user-generated text, paired with a compositional approach for irregular spellings, abbreviations, and code-mixing.

Cap4Bridge: Caption-Guided Cross-Modal Contextualization with Stochastic Augmentation for Text-Video Retrieval
IEEE Access 2026 MinJu Jeon, HyunGee Kim, Si-Woo Kim, Youngtaek Oh, Soeun Lee, Dong-Jin Kim 2026-03-01
#Text-Video Retrieval #Data-Centric

Bridges the text-video modality gap by using generated captions as cross-modal context, enriched through stochastic augmentation during training.

Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
EMNLP 2025 (Long, Main) MinJu Jeon, Si-Woo Kim, Ye-Chan Kim, HyunGee Kim, Dong-Jin Kim 2025-11-01
#Dense Video Captioning #Data-Centric

A dense video captioning framework that reweights video frames by saliency and adaptively retrieves relevant captions at inference time.

SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
ACM MM 2025 Si-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee, Taewhan Kim, Dong-Jin Kim 2025-10-01
#Zero-shot Captioning #Data-Centric

Refines noisy synthetic image-caption datasets through a one-to-many mapping that re-aligns each image with its best-matching captions.

© 2026 MinJu Jeon.
Built with Academic Portfolio Astro