Counterfactual Lesion Erasure for Measuring Visual Grounding in Medical Vision-Language Explanation

Authors

  • Yahya Ely Department of Applied Mathematics and Informatics, Higher Institute of Technological Education of Rosso, Route de Rosso, Rosso 16100, Mauritania Author
  • Bakar Sidaty Faculty of Legal and Economic Sciences, University of Aioun, Avenue de l’Indépendance, Aioun El Atrouss 51100, Mauritania Author
  • Hamadi Ahmed Department of Information Systems and Management, École Normale Supérieure de Nouakchott, Avenue Moktar Ould Daddah, Nouakchott 11200, Mauritania Author

Abstract

Medical vision-language systems often produce convincing explanatory sentences, but fluency alone does not show that the explanation is visually grounded. A model may name the correct abnormality while basing its wording on surrounding context, memorized report patterns, or language priors rather than the actual image region. This paper develops an empirical counterfactual lesion-erasure framework for testing whether generated medical explanations depend on the image evidence they claim to describe. We assembled 5,760 radiographic and dermatologic image-question pairs with expert lesion masks, generated short explanatory answers using five multimodal models, and then created counterfactual images by selectively erasing, preserving, or relocating clinically relevant regions. A grounding-sensitive model should reduce abnormality-specific explanation after lesion erasure, preserve it when irrelevant background is erased, and avoid transferring the original explanation to anatomically inconsistent relocated evidence. We introduce three quantitative measures: counterfactual explanation collapse, region-preservation stability, and relocation contradiction rate. We also estimate a latent grounding coefficient using a structured matrix model linking mask overlap, explanation content, and answer correctness. Across all models, ordinary answer accuracy was 78.6%, but only 54.2% of correct explanations showed expected collapse after lesion erasure. The proposed grounding coefficient separated visually dependent explanations from language-driven explanations with an area under the receiver operating characteristic curve of 0.817 against expert labels. The results indicate that medical vision-language evaluation should directly perturb claimed evidence regions rather than infer grounding from correctness or attention maps alone.

Downloads

Published

2025-12-04

How to Cite

Counterfactual Lesion Erasure for Measuring Visual Grounding in Medical Vision-Language Explanation. (2025). Studies in Data-Centric Computing, Cloud Frameworks, and Computational Intelligence, 15(12), 1-23. https://edgescholar.com/index.php/SDCCCI/article/view/e-2025-12-04