TY - JOUR
T1 - A comprehensive review of the video-to-text problem
AU - Perez-Martin, Jesus
AU - Bustos, Benjamin
AU - Guimarães, Silvio Jamil F.
AU - Sipiran, Ivan
AU - Pérez, Jorge
AU - Said, Grethel Coello
N1 - Publisher Copyright:
© 2021, The Author(s), under exclusive licence to Springer Nature B.V.
PY - 2022/6
Y1 - 2022/6
N2 - Research in the Vision and Language area encompasses challenging topics that seek to connect visual and textual information. When the visual information is related to videos, this takes us into Video-Text Research, which includes several challenging tasks such as video question answering, video summarization with natural language, and video-to-text and text-to-video conversion. This paper reviews the video-to-text problem, in which the goal is to associate an input video with its textual description. This association can be mainly made by retrieving the most relevant descriptions from a corpus or generating a new one given a context video. These two ways represent essential tasks for Computer Vision and Natural Language Processing communities, called text retrieval from video task and video captioning/description task. These two tasks are substantially more complex than predicting or retrieving a single sentence from an image. The spatiotemporal information present in videos introduces diversity and complexity regarding the visual content and the structure of associated language descriptions. This review categorizes and describes the state-of-the-art techniques for the video-to-text problem. It covers the main video-to-text methods and the ways to evaluate their performance. We analyze twenty-six benchmark datasets, showing their drawbacks and strengths for the problem requirements. We also show the progress that researchers have made on each dataset, we cover the challenges in the field, and we discuss future research directions.
AB - Research in the Vision and Language area encompasses challenging topics that seek to connect visual and textual information. When the visual information is related to videos, this takes us into Video-Text Research, which includes several challenging tasks such as video question answering, video summarization with natural language, and video-to-text and text-to-video conversion. This paper reviews the video-to-text problem, in which the goal is to associate an input video with its textual description. This association can be mainly made by retrieving the most relevant descriptions from a corpus or generating a new one given a context video. These two ways represent essential tasks for Computer Vision and Natural Language Processing communities, called text retrieval from video task and video captioning/description task. These two tasks are substantially more complex than predicting or retrieving a single sentence from an image. The spatiotemporal information present in videos introduces diversity and complexity regarding the visual content and the structure of associated language descriptions. This review categorizes and describes the state-of-the-art techniques for the video-to-text problem. It covers the main video-to-text methods and the ways to evaluate their performance. We analyze twenty-six benchmark datasets, showing their drawbacks and strengths for the problem requirements. We also show the progress that researchers have made on each dataset, we cover the challenges in the field, and we discuss future research directions.
KW - Deep learning
KW - Joint multi-modal embedding
KW - Matching-and-ranking
KW - Video captioning
KW - Video description retrieval
KW - Video-to-text
KW - Vision-and-language
KW - Visual-semantic embedding
KW - Visual-syntactic embedding
UR - http://www.scopus.com/inward/record.url?scp=85123167163&partnerID=8YFLogxK
U2 - 10.1007/s10462-021-10104-1
DO - 10.1007/s10462-021-10104-1
M3 - Article
AN - SCOPUS:85123167163
SN - 0269-2821
VL - 55
SP - 4165
EP - 4239
JO - Artificial Intelligence Review
JF - Artificial Intelligence Review
IS - 5
ER -