SOURCE-LINKED INTELLIGENCE
Organization of Valence and Arousal in Vision-Language Representations of Built Environments: Insights from the EMOIS Dataset
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of Spaces (EMOIS) dataset, comprising 1,544 real-world built-environment images. Each image is annotated with image-evoked valence and arousal ratings collected from Japanese adults by conducting a large-scale web-based survey, with approximately 120 ratings per image. Using Contra
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-06T23:12:30.000Z
First collected: 2026-09-20T21:12:06.801Z. This is not the publication date.