Deterministic Audit Trails for Reproducible Medical Vision-Language Model Evaluation Under Runtime and Interface Variability
- Authors
-
-
Narit Chaiyasit
Faculty of Information Technology, Nakhon Phanom University, 103 Moo 3, Ban Kaeng Subdistrict, Mueang Nakhon Phanom District, Nakhon Phanom 48000, ThailandAuthor -
Kittipong Rattanakorn
Department of Computer Science, Chiang Rai Rajabhat University, 80 Moo 9, Ban Du Subdistrict, Mueang Chiang Rai District, Chiang Rai 57100, ThailandAuthor
-
- Abstract
-
Medical vision-language model evaluation increasingly depends on complex software environments, external inference interfaces, preprocessing scripts, prompt wrappers, and automatic scoring components. Even when the clinical dataset is fixed, small differences in image conversion, runtime libraries, request serialization, model endpoint behavior, or output parsing can alter measured performance. Such variation complicates comparison between laboratories and weakens confidence in reported benchmark results. This paper presents an empirical study of deterministic audit trails for reproducible evaluation of medical vision-language models. We designed a trace-locked evaluation protocol in which each image-question instance is assigned a cryptographic content identifier, preprocessing fingerprint, prompt fingerprint, inference manifest, response digest, and scoring record. The protocol was tested on 9,600 medical image-question cases across radiography, computed tomography snapshots, ultrasound frames, and dermatology images. Four model interfaces, three runtime environments, and two scoring engines were evaluated under repeated runs. Without trace locking, nominally identical evaluations differed by 3.8 percentage points in strict answer accuracy and by 0.117 in macro calibration error across environments. With trace-locked preprocessing, deterministic request packaging, response canonicalization, and scorer version pinning, accuracy disagreement fell to 0.6 percentage points and calibration disagreement to 0.021. A discrepancy classifier trained on audit features identified 87.4% of nonreproducible cases before manual review. The findings show that evaluation reproducibility should be measured as a first-order property in medical vision-language benchmarking, particularly when multiple institutions, software stacks, or hosted model endpoints are involved.
- Downloads
- Published
- 2026-01-04
- Section
- Articles