Open library

This portal has been archived. Explore the next generation of this technology.

A Revised Generative Evaluation of Visual Dialogue

lib:7c0bd5d18f2c8e55 (v1.0.0)

Authors: Daniela Massiceti,Viveka Kulharia,Puneet K. Dokania,N. Siddharth,Philip H. S. Torr
ArXiv: 2004.09272
Document: PDF DOI

Abstract URL: https://arxiv.org/abs/2004.09272v2

Evaluating Visual Dialogue, the task of answering a sequence of questions relating to a visual input, remains an open research challenge. The current evaluation scheme of the VisDial dataset computes the ranks of ground-truth answers in predefined candidate sets, which Massiceti et al. (2018) show can be susceptible to the exploitation of dataset biases. This scheme also does little to account for the different ways of expressing the same answer--an aspect of language that has been well studied in NLP. We propose a revised evaluation scheme for the VisDial dataset leveraging metrics from the NLP literature to measure consensus between answers generated by the model and a set of relevant answers. We construct these relevant answer sets using a simple and effective semi-supervised method based on correlation, which allows us to automatically extend and scale sparse relevance annotations from humans to the entire dataset. We release these sets and code for the revised evaluation scheme as DenseVisDial, and intend them to be an improvement to the dataset in the face of its existing constraints and design choices.

Relevant initiatives

Related knowledge about this paper

Search on this portal

Reproduced results (crowd-benchmarking and competitions)

Artifact and reproducibility checklists

Common formats for research projects and shared artifacts

Collective Knowledge (organizing research projects based on FAIR principles)

Reproducibility initiatives

Comments

Please log in to add your comments!

If you notice any inapropriate content that should not be here, please report us as soon as possible and we will try to remove it within 48 hours!

A Revised Generative Evaluation of Visual Dialogue

Relevant initiatives Hide

Comments Hide

Relevant initiatives

Comments