ICASSP 2026 (Prof. Chanho Eom, 1 paper)
Perceptual Artificial Intelligence (Perceptual AI Lab) (Gyuwon Han*, Young Kyun Jang*, Chanho Eom)
We are delighted to announce that one paper from the Perceuptual AI Lab (PAI Lab, Prof. Chanho Eom) has been accepted to the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).
Title:
COVA: Text-Guided Composed Retrieval for Audio-Visual Content
Authors:
Gyuwon Han*, Young Kyun Jang*, Chanho Eom
Abstract:
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing benchmarks consider only visual changes, ignoring videos that differ in audio despite visual similarity. To address this limitation, we introduce Composed retrieval for Video with its Audio (COVA), a new retrieval task that accounts for both visual and auditory variations. To support this, we construct AV-Comp, a benchmark of video pairs with cross-modal changes and textual queries describing the differences, enabling retrieval based on audio as well. We also propose AVT Compositional Fusion (AVT), which integrates video, audio, and text features by selectively aligning the query to the most relevant modality. AVT outperforms traditional unimodal fusion and serves as a strong baseline for COVA.
| 이전글 | ICASSP 2026 (Prof. Jihun Kim, 2 papers) |
|---|---|
| 다음글 | ICLR 2026 (Prof. Jihyong Oh, 1 paper) |