Spatiotemporal localization of activities in untrimmed surveillance videos is a hard task, especially given the occurrence of simultaneous activities across different temporal and spatial scales. We tackle this problem using a cascaded region proposal and detection (CRPAD) framework implementing frame-level simultaneous action detection, followed by tracking. We propose the use of a frame-level spatial detection model based on advances in object detection and a temporal linking algorithm that models the temporal dynamics of the detected activities. We show results on the VIRAT dataset through the recent Activities in Extended Video (ActEV) challenge that is part of the TrecVID competition[1, 2].
Guangle YaoTao LeíXianyuan LiuPing Jiang
Moon YoungHyung-Il KimJongyoul Park
Zheng ShouJunting PanJ. ChanKazuyuki MiyazawaHassan MansourAnthony VetroXavier Giró-i-NietoShih‐Fu Chang