A Latent Diffusion Framework for Generating Counterfactual Radiographs to Disentangle Visual Grounding from Lexical Invariance in Multimodal Diagnosis
- Authors
-
-
Hoang Nguyen
Faculty of Information Technology, Quy Nhon University, 170 An Duong Vuong, Quy Nhon, Binh Dinh, VietnamAuthor -
Bich Tran
Faculty of Computer Science and Engineering, Thuyloi University, 175 Tay Son, Dong Da, Hanoi, VietnamAuthor -
Cuong Le
Department of Diagnostic Imaging, Hue University of Medicine and Pharmacy, 06 Ngo Quyen, Hue, Thua Thien Hue, VietnamAuthor
-
- Abstract
-
Multimodal foundation models deployed for medical visual question answering frequently rely on superficial linguistic correlations rather than genuine anatomical visual evidence. When clinical questions undergo syntactic paraphrasing, contemporary architectures alter diagnostic conclusions in over thirty percent of cases, revealing severe epistemic fragility. However, auditing whether these diagnostic flips stem from visual grounding failure or language prior bias remains fundamentally constrained by observational medical datasets, where pathological findings cannot be causally manipulated independently of patient anatomical confounders. To resolve this empirical bottleneck, this investigation introduces a mask-guided latent diffusion framework designed to synthesize high-fidelity counterfactual chest radiographs. Our architecture formulates a localized reverse stochastic differential equation trajectory modulated by cross-attention pathology steering and structural structural-similarity regularization, enabling precise surgical insertion and removal of focal thoracic lesions while strictly preserving underlying patient identity and bony morphology. Using this generative engine, we construct a certified benchmark suite comprising 4,000 paired observational-counterfactual radiographs across pneumothorax, pleural effusion, and cardiomegaly. Multi-reader clinical Turing audits by board-certified radiologists demonstrate 94.2% structural realism, confirming that synthetic modifications remain indistinguishable from authentic physiological presentations. Systematic stress-testing of nine prominent vision-language models across factual and counterfactual pairs demonstrates that language shortcuts account for 68.4% of observed paraphrase flips. Counterfactual preference tuning reduces prompt sensitivity by 84.1% while elevating true visual grounding alignment, providing a reproducible audit protocol for pre-market algorithmic safety certification.
- Downloads
- Published
- 2026-06-04
- Section
- Articles