文章基本信息

标题：A Video Classification Method Based on Spatiotemporal Detail Attention and Feature Fusion
本地全文：下载
作者：Xuchao Gong ; Zongmin Li
期刊名称：Mobile Information Systems
印刷版ISSN：1574-017X
出版年度：2022
卷号：2022
DOI：10.1155/2022/4213335
语种：English
出版社：Hindawi Publishing Corporation
摘要：With the explosive growth of Internet video data, demands for accurate large-scale video classification and management are increasing. In the real-world deployment, the balance between effectiveness and timeliness should be fully considered. Generally, the video classification algorithm equipped with time segment network is used in industrial deployment, and the frame extraction feature is used to classify video actions However, the issue of semantic deviation will be raised due to coarse feature description. In this paper, we propose a novel method, called image dense feature and internal significant detail description, to enhance the generalization and discrimination of feature description. Specifically, the location information layer of space-time geometric relationship is added to effectively engrave the local features of convolution layer. Moreover, the multimodal feature graph network is introduced to effectively improve the generalization ability of feature fusion. Extensive experiments show that the proposed method can effectively improve the results on two commonly used benchmarks (kinetics 400 and kinetics 600).