Video Visual Relation Detection
Patent Name: Video Visual Relation Detection
Filing No: 62/546,641 (US)
Filing Date: 17 Aug 2017
Country to be Filed: USA
Description: As a bridge to connect vision and language, visual relations between objects in the form of relation triplet hsubject,predicate, objecti, such as “person-touch-dog” and “cat-above-sofa”, provide a more comprehensive visual content understanding beyond objects. In this paper, we propose a novel vision task named Video Visual Relation Detection (VidVRD) to perform visual relation detection in videos instead of still images (ImgVRD). As compared to still images, videos provide a more natural set of features for detecting visual relations, such as the dynamic relations like “A-follow-B” and “A-towards-B”, and temporally changing relations like “Achase-B” followed by “A-hold-B”. However, VidVRD is technically more challenging than ImgVRD due to the diculties in accurate object tracking and diverse relation appearances in video domain. To this end, we propose a VidVRD method, which consists of object tracklet proposal, short-term relation prediction and greedy relational association. Moreover, we contribute the rst dataset for VidVRD evaluation, which contains 1,000 videos with manually labeled visual relations, to validate our proposed method. On this dataset, our method achieves the best performance in comparison with the state-of-the-art baselines.