Abstract
Image captioning requires integrating accurate object recognition with a deep understanding of semantic relationships to generate natural descriptions. Existing methods primarily rely on grid-based visual features for image understanding, which suffer from semantic fragmentation and fail to form coherent region-level representations. Furthermore, many approaches do not effectively integrate semantic prior information from the textual modality,
resulting in coarse and inaccurate descriptions. To address these challenges, we propose the Semantic Merging Adaptive Regional Transformer (SMART) model, which incorporates three novel modules: a Spatial Continuous Region Aggregation Module (SCRAM) to mitigate fragmentation by aggregating grid
features into complete object regions; a Semantic-aware Region Enhancement Module (SREM) to align textual semantic priors with these region representations; and an Adaptive Region-to-Grid Feature Fusion Module (ARGFFM) that dynamically selects key regions to enhance feature capacity. Finally, these enhanced features are subsequently leveraged by a transformer-based encoder-decoder architecture to generate semantically robust and accurate captions. Extensive experimental results on the MS COCO dataset establish that the SMART model achieves state of-the-art performance, confirming the efficacy of the SMART.
resulting in coarse and inaccurate descriptions. To address these challenges, we propose the Semantic Merging Adaptive Regional Transformer (SMART) model, which incorporates three novel modules: a Spatial Continuous Region Aggregation Module (SCRAM) to mitigate fragmentation by aggregating grid
features into complete object regions; a Semantic-aware Region Enhancement Module (SREM) to align textual semantic priors with these region representations; and an Adaptive Region-to-Grid Feature Fusion Module (ARGFFM) that dynamically selects key regions to enhance feature capacity. Finally, these enhanced features are subsequently leveraged by a transformer-based encoder-decoder architecture to generate semantically robust and accurate captions. Extensive experimental results on the MS COCO dataset establish that the SMART model achieves state of-the-art performance, confirming the efficacy of the SMART.
| Original language | English |
|---|---|
| Number of pages | 11 |
| Journal | IEEE Transactions on Multimedia |
| Early online date | 27 Jul 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 27 Jul 2026 |
Funding
This work was supported by China Scholarship Council, the Major Project of the Anhui Provincial University Scientific Research Program (2025AHGXZK20063), the Natural Science Foundation of the Education Bureau of Anhui Province (2024AH051583), the open project of Key Laboratory of Intelligent Computing & Signal Processing, Ministry of Education (Anhui University) (2025002).
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
-
SDG 12 Responsible Consumption and Production
Keywords
- Image Captioning
- Semantic Region Representation
- Semantic Prior Information
- Cross-modal Alignment
Fingerprint
Dive into the research topics of 'SMART: Semantic Merging Adaptive Regional Transformer for Image Captioning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver