×

Article

Image Captioning Using OpenCV, CNN and LSTM

Author(s) : Yegireddi Ramesh

  • Abstract
  • Download PDF (9)
  • XML
  • Views (11)
  • Video
  • Audio
  • Citation
  • Certificate

Image captioning is one of the most important domains in human-computer interaction that generates a description of the visual information. In this study, the need for large-scale studies is motivated to investigate the benefits of deep learning techniques including the combination of CNN and LSTM models on image captioning. The proposed approach extracts visual features by CNN from the diverse images and generates coherent and contextually appropriate captions using the LSTM model from the Flickr8K database, which consists of various images with respective captions. The combination of CNN and LSTM layers is used to capture the spatial and temporal features, which improves the accuracy and relevance of the generated text descriptions. In the proposed method, the combination of CNN and LSTM layers allows for the efficient extraction of spatial and temporal information, significantly improving the accuracy and relevance of generated text descriptions. The purpose of this study is to measure the performance of CNN-LSTM model for image captioning, with the aim of enhancing visual understanding and boosting developments in natural language processing.

Readers Around the World
Similar Papers
Login Register