Skip to main navigation Skip to search Skip to main content

SMART: Semantic Merging Adaptive Regional Transformer for Image Captioning

  • Fengling Jiang
  • , Yujin Cao
  • , Le Zou*
  • , Erfu Yang
  • , Chenglong Li
  • , Chaomin Luo
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

1 Downloads (Pure)

Abstract

Image captioning requires integrating accurate object recognition with a deep understanding of semantic relationships to generate natural descriptions. Existing methods primarily rely on grid-based visual features for image understanding, which suffer from semantic fragmentation and fail to form coherent region-level representations. Furthermore, many approaches do not effectively integrate semantic prior information from the textual modality,
resulting in coarse and inaccurate descriptions. To address these challenges, we propose the Semantic Merging Adaptive Regional Transformer (SMART) model, which incorporates three novel modules: a Spatial Continuous Region Aggregation Module (SCRAM) to mitigate fragmentation by aggregating grid
features into complete object regions; a Semantic-aware Region Enhancement Module (SREM) to align textual semantic priors with these region representations; and an Adaptive Region-to-Grid Feature Fusion Module (ARGFFM) that dynamically selects key regions to enhance feature capacity. Finally, these enhanced features are subsequently leveraged by a transformer-based encoder-decoder architecture to generate semantically robust and accurate captions. Extensive experimental results on the MS COCO dataset establish that the SMART model achieves state of-the-art performance, confirming the efficacy of the SMART.
Original languageEnglish
Number of pages11
JournalIEEE Transactions on Multimedia
Early online date27 Jul 2026
DOIs
Publication statusE-pub ahead of print - 27 Jul 2026

Funding

This work was supported by China Scholarship Council, the Major Project of the Anhui Provincial University Scientific Research Program (2025AHGXZK20063), the Natural Science Foundation of the Education Bureau of Anhui Province (2024AH051583), the open project of Key Laboratory of Intelligent Computing & Signal Processing, Ministry of Education (Anhui University) (2025002).

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure
  2. SDG 12 - Responsible Consumption and Production
    SDG 12 Responsible Consumption and Production

Keywords

  • Image Captioning
  • Semantic Region Representation
  • Semantic Prior Information
  • Cross-modal Alignment

Fingerprint

Dive into the research topics of 'SMART: Semantic Merging Adaptive Regional Transformer for Image Captioning'. Together they form a unique fingerprint.

Cite this