Computer Science Faculty Publications
Who Matters More in Radiology Report Generation: Vision Encoders or Language Models?
Document Type
Conference Proceeding
Publication Date
3-10-2026
Abstract
The rapid development of Multimodal Large Language Models (MLLMs) has advanced Radiology Report Generation (RRG). While much of this progress is driven by increasingly powerful Large Language Models (LLMs), the roles of both the vision encoder and the LLM remain underexplored, especially in domain-specific contexts. In this work, we systematically study how different vision encoders and LLMs affect RRG performance, analyzing the task from both vision- and languagecentric perspectives. Through extensive evaluation, we show that domain-adapted vision encoders and LLMs significantly enhance the quality and clinical relevance of generated reports. These findings offer practical guidance for building effective MLLMs in medical imaging.
Recommended Citation
Zhao, Kun, Yang Du, Rhianna Zhang, Liang Zhan, Dongkuan Xu, Pengfei Gu, and Haoteng Tang. "Who Matters More in Radiology Report Generation: Vision Encoders or Language Models?." In 2025 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 1985-1989. IEEE, 2025. https://doi.org/10.1109/ICDMW69685.2025.00239
Publication Title
2025 IEEE International Conference on Data Mining Workshops (ICDMW)
DOI
10.1109/ICDMW69685.2025.00239

Comments
©2025 IEEE