Cheddi, F., Habbani, A. and Nait-Charif, H., 2024. Improving the CXR reports generation with multi-modal feature alignment and self-refining strategy. In: 2024 3rd International Conference on Embedded Systems and Artificial Intelligence (ESAI), 19-20 December 2024, Fez, Morocco.
Full text available as:
Preview |
PDF
Improving the CXR Reports.pdf - Accepted Version 557kB |
|
Copyright to original material in this document is with the original owner(s). Access to this content through BURO is granted on condition that you use it only for research, scholarly or other non-commercial purposes. If you wish to use it for any other purposes, you must contact BU via BURO@bournemouth.ac.uk. Any third party copyright material in this document remains the property of its respective owner(s). BU grants no licence for further use of that third party material. |
Official URL: https://ieeexplore.ieee.org/xpl/conhome/10913474/p...
DOI: 10.1109/ESAI62891.2024.10913509
Abstract
Medical chest X-ray images are necessary for the diagnosis of different diseases. The need for automated interpretation and report generation of these images is crucial. It not only saves radiologists' time but also minimizes the risk of diagnostic mistakes. However, several challenges impede this task due to the employment of uni-directional image-To-report in the encoder-decoder deep learning model and the absence of contextual details in the process of report generation can lead to incomplete or inaccurate descriptions. To address this, we propose an approach based on Multi-modal feature Alignment and Self-Refining mechanism RG-MASR in order to generate an improved medical report from chest X-ray images automatically. Our method comprises three modules: visual and textual characteristics extraction to extract the semantic characteristics from X-ray images and their paired reports. Second, a multimodal feature alignment module is employed to leverage both textual and visual features. Finally, we integrate a self-refining technique in the report generator module to refine alignment and improve the output to generate a comprehensive report. We evaluate our method on the IU X-ray and NIH public chest X-ray datasets. The results demonstrate that our proposed RG-MASR surpasses existing approaches in terms of ROUGE and BLEU metrics.
| Item Type: | Conference or Workshop Item (Paper) |
|---|---|
| Uncontrolled Keywords: | CXR report generation; Multi-modal alignment; medical image; Deep learning; Medical report; Transformer |
| Group: | Faculty of Media, Science and Technology |
| ID Code: | 42166 |
| Deposited By: | Symplectic RT2 |
| Deposited On: | 02 Sep 2026 15:42 |
| Last Modified: | 02 Sep 2026 15:42 |
Downloads
Downloads per month over past year
| Repository Staff Only - |
Tools
Tools