EventsMOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed.
Published
This submission belongs to the session 03. USEDAT-02: USA-Europe Data Analysis Training Program Workshop, Cambridge, UK-Bilbao, Spain-Miami, USA, 2016 of the event MOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed.
Published date
30 Dec, 2016
Citation
Xin Xu, Haibin Liu, Yi Ji, Xin Lin, Chunping Liu, Video Description with Spatio-temporal Feature and Knowledge Transferring, in Proceedings of MOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed., 15 October–20 October 2022, MDPI: Basel, Switzerland, doi: 10.3390/mol2net-02-03851
Share
Email
Facebook
Twitter
LinkedIn

Video Description with Spatio-temporal Feature and Knowledge Transferring

Haibin Liu 1
Yi Ji 1
Xin Lin 1
1. School of Computer Science and Technology, Soochow University
2. Collaborative Innovation Center of Novel Software Technology and Industrialization
3. Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education
Abstract

Describing open-domain video with natural language sequence is a major challenge for computer vision. In this paper, we investigate how to use temporal information and learn linguistic knowledge for video description. Traditional convolutional neural networks (CNN) can only learn powerful spatial features in the videos, but they ignored underlying temporal features. To solve this problem, we extract SIFT flow features to get temporal information. Sequence generator of recent work are solely trained on text from video description datasets, so the sequence generated tend to show linguistic irregularities associated with a restricted language model and small vocabulary. For this, we transfer knowledge from large text corpora and employ word2vec to be the word representation. The experimental results have demonstrated that our model outperforms related work.

Keywords
video description
SIFT flow
knowledge transferring
word2vec
Poster
abstract.pdf
Building Domain-Specific Sentiment Lexicon by Sentiment Seed Expansion
Fusing Augmented Spatio-temporal Features for Action Recognition