Publications & Projects

Publications & Projects

ICASSP 2026 (Prof. Chanho Eom, 1 paper)
  • Title

    ICASSP 2026 (Prof. Chanho Eom, 1 paper)

  • Authors

    Perceptual Artificial Intelligence (Perceptual AI Lab) (Gyuwon Han*, Young Kyun Jang*, Chanho Eom)

  • Audio-Visual Retrieval
  • Multimodal Fusion
  • Composed Retrieval
Abstract

We are delighted to announce that one paper from the Perceuptual AI Lab (PAI Lab, Prof. Chanho Eom) has been accepted to the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).

Title:
COVA: Text-Guided Composed Retrieval for Audio-Visual Content

Authors:
Gyuwon Han*, Young Kyun Jang*, Chanho Eom

Abstract:
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing benchmarks consider only visual changes, ignoring videos that differ in audio despite visual similarity. To address this limitation, we introduce Composed retrieval for Video with its Audio (COVA), a new retrieval task that accounts for both visual and auditory variations. To support this, we construct AV-Comp, a benchmark of video pairs with cross-modal changes and textual queries describing the differences, enabling retrieval based on audio as well. We also propose AVT Compositional Fusion (AVT), which integrates video, audio, and text features by selectively aligning the query to the most relevant modality. AVT outperforms traditional unimodal fusion and serves as a strong baseline for COVA.

이전글 ICASSP 2026 (Prof. Jihun Kim, 2 papers)
다음글 ICLR 2026 (Prof. Jihyong Oh, 1 paper)