Learning to Compose and Reason with Language Tree Structures for Visual Grounding

IEEE Transactions on Pattern Analysis and Machine Intelligence(2022)

引用 119|浏览485
暂无评分
摘要
Grounding natural language in images, such as localizing “the black dog on the left of the tree”, is one of the core problems in artificial intelligence, as it needs to comprehend the fine-grained language compositions. However, existing solutions merely rely on the association between the holistic language features and visual features, while neglect the nature of composite reasoning implied in th...
更多
查看译文
关键词
Grounding,Visualization,Dogs,Natural languages,Cognition,Computational modeling,Semantics
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要