AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Learning to Refer from Estimated Listener Gaze

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are sampled from the speaker policy being optimized, conditioned on images and target referents; then, a neural listener estimating human gaze behavior maps from images and sampled referring expressions to scanpaths, each represented by a sequence of fixations, with each fixation corresponding to a word in the referring expression

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:41:04.663Z. This is not the publication date.