We are pleased to share that our latest research paper titled “DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval” has been submitted to Pattern Recognition and is also available as a preprint on arXiv.
This work introduces DREAM (Dual-path Representation Enhancement and Alignment Model), a novel vision-language framework for text-to-video retrieval. DREAM combines a dual-objective language encoding strategy with a hierarchical vision encoder incorporating cascaded group attention to effectively model complex visual and textual relationships. Extensive experiments on the benchmark MSRVTT, MSVD, and LSMDC datasets demonstrate state-of-the-art retrieval performance, highlighting the effectiveness of the proposed architecture for robust and context-aware cross-modal representation learning.
