EventsMOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed.
Published
This submission belongs to the session 03. USEDAT-02: USA-Europe Data Analysis Training Program Workshop, Cambridge, UK-Bilbao, Spain-Miami, USA, 2016 of the event MOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed.
Published date
30 Dec, 2016
Citation
Rui Ge, Xiaoyi Wan, Yi Ji, Chunping Liu, Shengrong Gong, Fusing Augmented Spatio-temporal Features for Action Recognition, in Proceedings of MOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed., 15 October–20 October 2022, MDPI: Basel, Switzerland, doi: 10.3390/mol2net-02-03852
Share
Email
Facebook
Twitter
LinkedIn

Fusing Augmented Spatio-temporal Features for Action Recognition

Xiaoyi Wan 1
Yi Ji 1
1. School of Computer Science and Technology, Soochow University
2. Collaborative Innovation Center of Novel Software Technology and Industrialization
3. Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University
4. School of Computer Science and Engineering, Changshu Institute of Technology
Abstract

Visual features are vitally important for action recognition in videos. However, traditional features fail to effectively recognize actions for two reasons: on one hand, spatial features are not powerful enough to capture appearance information of complex video actions; on the other hand, important temporal details are always ignored when pooling and encoding. In this paper, we present a new architecture that fuses multiple augmented spatio-temporal features. In order to strengthen spatial features, we conduct crop and horizontal flip on original frame images. Then we feed these processed images into deep Two-Stream network to produce robust spatial representations. To get powerful temporal features, we employ fourier temporal pyramid (FTP) to capture three different levels of video context, including short-term level, medium-range level, and global-range level. At last, we fuse these augmented spatio-temporal features using canonical correlation analysis (CCA) method, which is capable to capture the correlation between these features. Experimental results on UCF101 dataset show that our method can achieve excellent performance for action recognition.

Keywords
action recognition
CNN features
fourier temporal pyramid
CCA fusion
Poster
paper.pdf
Video Description with Spatio-temporal Feature and Knowledge Transferring
Attention-based CNNs for Aspect-level Sentiment Classification