Dr. Marc A. Kastner

About me

Other languages

Deutsch

Esperanto

日本語

Estimating the visual variety of concepts by referring to Web popularity

Back to publications

Authors: Marc A. Kastner, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase

Abstract:

Increasingly sophisticated methods for data processing demand knowledge on the semantic relationship between language and vision. New fields of research like Explainable AI demand to step away from black-boxed approaches and understanding how the underlying semantics of data sets and AI models work. Advancements in Psycholinguistics suggest, that there is a relationship from language perception to how language production and sentence creation work. In this paper, a method to measure the visual variety of concepts is proposed to quantify the semantic gap between vision and language. For this, an image corpus is recomposed using ImageNet and Web data. Web-based metrics for measuring the popularity of sub-concepts are used as a weighting to ensure that the image composition in a dataset is as natural as possible. Using clustering methods, a score describing the visual variety of each concept is determined. A crowd-sourced survey is conducted to create ground-truth values applicable for this research. The evaluations show that the recomposed image corpus largely improves the measured variety compared to previous datasets. The results are promising and give additional knowledge about the relationship of language and vision.

Type: Journal paper at Multimedia Tools and Applications (MTAP), 78(7), 9463-9488

Publication date: April 2019

DOI: 10.1007/s11042-018-6528-x

Attached Files

preprint

If you have questions or ideas about this research, feel free to leave a comment below or send me an email. I will reply quickly.