SOURCE-LINKED INTELLIGENCE
E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews
Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation.
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T10:07:26.000Z
- arXiv · Artificial Intelligence · 2026-09-17T10:07:26.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.