Dynamic Graph Attention for Referring Expression Comprehension

Yang, S; Li, G; Yu, Y

File Download

There are no files associated with this item.

Links for fulltext

(May Require Subscription)

Publisher Website: 10.1109/ICCV.2019.00474
Scopus: eid_2-s2.0-85081925645
WOS: WOS:000531438104079
Find via

Supplementary

Citations:
- Scopus: 0
- Web of Science: 0
Appears in Collections:
- Computer Science: Conference papers

Conference Paper: Dynamic Graph Attention for Referring Expression Comprehension

Title	Dynamic Graph Attention for Referring Expression Comprehension
Authors	Yang, S Li, G Yu, Y
Issue Date	2019
Publisher	Institute of Electrical and Electronics Engineers. The Journal's web site is located at http://ieeexplore.ieee.org/xpl/conhome.jsp?punumber=1000149
Citation	Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, 27 October - 2 November 2019, p. 4643-4652 How to Cite? DOI: http://dx.doi.org/10.1109/ICCV.2019.00474
Abstract	Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among the objects in the image. Meanwhile, the visual reasoning process is guided by the linguistic structure of the referring expression. However, existing approaches treat the objects in isolation or only explore the first-order relationships between objects without being aligned with the potential complexity of the expression. Thus it is hard for them to adapt to the grounding of complex referring expressions. In this paper, we explore the problem of referring expression comprehension from the perspective of language-driven visual reasoning, and propose a dynamic graph attention network to perform multi-step reasoning by modeling both the relationships among the objects in the image and the linguistic structure of the expression. In particular, we construct a graph for the image with the nodes and edges corresponding to the objects and their relationships respectively, propose a differential analyzer to predict a language-guided visual reasoning process, and perform stepwise reasoning on top of the graph to update the compound object representation at every node. Experimental results demonstrate that the proposed method can not only significantly surpass all existing state-of-the-art algorithms across three common benchmark datasets, but also generate interpretable visual evidences for stepwisely locating the objects referred to in complex language descriptions.
Persistent Identifier	http://hdl.handle.net/10722/284142
ISSN	1550-5499 2023 SCImago Journal Rankings: 12.263
ISI Accession Number ID	WOS:000531438104079

DC Field	Value	Language
dc.contributor.author	Yang, S	-
dc.contributor.author	Li, G	-
dc.contributor.author	Yu, Y	-
dc.date.accessioned	2020-07-20T05:56:25Z	-
dc.date.available	2020-07-20T05:56:25Z	-
dc.date.issued	2019	-
dc.identifier.citation	Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, 27 October - 2 November 2019, p. 4643-4652	-
dc.identifier.issn	1550-5499	-
dc.identifier.uri	http://hdl.handle.net/10722/284142	-
dc.description.abstract	Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among the objects in the image. Meanwhile, the visual reasoning process is guided by the linguistic structure of the referring expression. However, existing approaches treat the objects in isolation or only explore the first-order relationships between objects without being aligned with the potential complexity of the expression. Thus it is hard for them to adapt to the grounding of complex referring expressions. In this paper, we explore the problem of referring expression comprehension from the perspective of language-driven visual reasoning, and propose a dynamic graph attention network to perform multi-step reasoning by modeling both the relationships among the objects in the image and the linguistic structure of the expression. In particular, we construct a graph for the image with the nodes and edges corresponding to the objects and their relationships respectively, propose a differential analyzer to predict a language-guided visual reasoning process, and perform stepwise reasoning on top of the graph to update the compound object representation at every node. Experimental results demonstrate that the proposed method can not only significantly surpass all existing state-of-the-art algorithms across three common benchmark datasets, but also generate interpretable visual evidences for stepwisely locating the objects referred to in complex language descriptions.	-
dc.language	eng	-
dc.publisher	Institute of Electrical and Electronics Engineers. The Journal's web site is located at http://ieeexplore.ieee.org/xpl/conhome.jsp?punumber=1000149	-
dc.relation.ispartof	IEEE International Conference on Computer Vision (ICCV) Proceedings	-
dc.rights	IEEE International Conference on Computer Vision (ICCV) Proceedings. Copyright © Institute of Electrical and Electronics Engineers.	-
dc.rights	©2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.	-
dc.title	Dynamic Graph Attention for Referring Expression Comprehension	-
dc.type	Conference_Paper	-
dc.identifier.email	Yu, Y: yzyu@cs.hku.hk	-
dc.identifier.authority	Yu, Y=rp01415	-
dc.identifier.doi	10.1109/ICCV.2019.00474	-
dc.identifier.scopus	eid_2-s2.0-85081925645	-
dc.identifier.hkuros	310940	-
dc.identifier.spage	4643	-
dc.identifier.epage	4652	-
dc.identifier.isi	WOS:000531438104079	-
dc.publisher.place	United States	-
dc.identifier.issnl	1550-5499	-

File Download

Links for fulltext

(May Require Subscription)

Supplementary

Conference Paper: Dynamic Graph Attention for Referring Expression Comprehension

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats