登录    注册    忘记密码

详细信息

BPA-Net: Bidirectional Prototype Alignment for Image Captioning  ( EI收录)  

文献类型:期刊文献

英文题名:BPA-Net: Bidirectional Prototype Alignment for Image Captioning

作者:Li, Tianci[1]; Xu, Cheng[1]; Song, Yizhe[1]; Yuan, Jiazheng[2]; Dai, Songyin[1]; Zhang, Jiancheng[1]

第一作者:Li, Tianci

机构:[1] Beijing Union University, Beijing Key Laboratory of Information Service Engineering, Beijing, 100101, China; [2] Capital Normal University, Research Center For Language Intelligence of China, Beijing, 100048, China

第一机构:北京联合大学北京市信息服务工程重点实验室

通讯机构:[1]Beijing Union University, Beijing Key Laboratory of Information Service Engineering, Beijing, 100101, China|[11417103]北京联合大学北京市信息服务工程重点实验室;[11417]北京联合大学;

年份:2026

外文期刊名:IEEE Transactions on Circuits and Systems for Video Technology

收录:EI(收录号:20262520929802);Scopus(收录号:2-s2.0-105041933547)

语种:英文

外文关键词:Alignment - Semantics - Software prototyping

摘要:Image captioning aims to generate accurate textual descriptions for visual inputs. A major challenge is the semantic gap between continuous visual features and discrete linguistic concepts. Existing methods often rely on the unidirectional injection of external textual priors, which inevitably introduces semantic noise. Furthermore, CLIP-based image captioning methods lack direct supervision from ground-truth captions, leading to weak cross-modal alignment. To address these issues, we propose BPA-Net, a framework for bidirectional prototype alignment. Specifically, the Dual Prototype Extraction module independently projects visual features and ground-truth text tokens into a shared semantic space to form structured prototypes. Then, the Cross-Modal Prototype Alignment module enables bidirectional interaction and alignment between visual and textual prototypes guided by symmetric matching objectives. This approach effectively mitigates the noise inherent in external proxy texts and establishes robust semantic alignment. Extensive experiments on the MS-COCO, Flickr8k, and Flickr30k benchmarks demonstrate that BPA-Net achieves state-of-the-art performance. Furthermore, evaluations on the Nocaps dataset confirm its strong generalization capability for describing novel objects. ? 1991-2012 IEEE.

参考文献:

正在载入数据...

版权所有©北京联合大学 重庆维普资讯有限公司 渝B2-20050021-8 
渝公网安备 50019002500408号 违法和不良信息举报中心