AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Beyond OCR Accuracy: Text-Centric VQA Under Image Degradation with Modular and End-to-End

arXiv · AI, language, vision and robotics · article · Sep 12, 2026 · UTC

Text-centric Visual Question Answering (VQA) requires reading and reasoning over text embedded in images, a task made substantially harder when images suffer from real-world degradation such as motion blur, low resolution, or compression artifacts. While modular OCR-based pipelines and end-to-end vision-language models are both widely used for this task, their comparative robustness under degraded conditions remains underexplored. We present an empirical study comparing two modular pipelines with SA-DBNet, a custom detector architecture combining ResNet-18 with self-attention spatial modeling

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T16:41:15.630Z. This is not the publication date.