Difference Visual Question Answering in Medical Imaging Based on Multimodal Large Models and Enhanced by Prior Region-level Difference Descriptions

Authors

  • Bokai Yang Institute of Intelligence Science and Engineering, Shenzhen Polytechnic University, Shenzhen 518055, China; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China
  • Haorong Li Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China
  • Yirong Qin School of Urban Governance and Public Affairs, Suzhou City University, Suzhou 215000, China
  • Huazhen Huang Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China
  • Ye Li Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China
  • Jikui Liu Institute of Intelligence Science and Engineering, Shenzhen Polytechnic University, Shenzhen 518055, China
  • Yunpeng Cai Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China

DOI:

https://doi.org/10.5755/j01.itc.55.2.44202

Keywords:

Large language models (LLMs), Multimodal learning, Visual question answering (VQA), Medical imaging, Difference visual question answering (Diff-VQA), Radiology report generation

Abstract

Difference Visual Question Answering (Diff-VQA) in medical imaging automatically compares patient images across time points to support assessment of lesion progression and treatment efficacy. However, pixel-level matching is unreliable due to non-rigid deformations, view shifts, and acquisition noise, while existing models often rely on synthetic labels and lack effective integration of local and global information. To address these challenges, we propose a multimodal large-model framework that adopts a progressive “local semantic modeling–global difference reasoning” strategy. Key anatomical regions in chest X-rays are localized via object detection and aligned with VinDr-CXR annotations to construct region–disease mappings, transforming misalignment into semantic difference analysis. A dynamic sampling strategy further generates clinically meaningful image pairs with fine-grained difference labels. Finally, a multimodal large model fuses local features with global context to support single-image QA, dual-image disease description, and global difference reasoning. Experiments on the MIMIC-Diff-VQA dataset demonstrate state-of-the-art accuracy in single-image QA and substantial improvements in Diff-VQA tasks over mainstream medical large models. In the single-image QA tasks, our model improves accuracy from 52.5% to 64.1% (22.2% relative improvement), and in the Diff-VQA tasks, the CIDEr score increases from 1.027 to 1.379 (34.3% relative improvement). These results highlight the framework’s potential to enhance diagnostic accuracy and strengthen clinical decision support in radiology practice.

Downloads

Published

2026-07-23

Issue

Section

Articles