Abstract
This paper presents a novel and comprehensive framework for video anomaly detection, distinguished by its specialized spatio-temporal feature extraction and precise anomaly prediction capabilities. The proposed system employs an advanced spatio-temporal attention-based framework designed for effective video frame reconstruction. By isolating and amplifying critical feature regions within the frames, it enables the extraction of fine-grained spatial and temporal representations, which are crucial for detecting subtle anomalies. Complementing this, an attentive U-Net architecture is employed to predict anomalies with high precision, incorporating motion features to enhance temporal coherence and anomaly localization. The attention mechanism in both components is strategically designed to focus on critical areas within each frame and sequence, where abnormal activities are likely to occur, improving detection accuracy and reducing false positives. The two components are seamlessly integrated using a fusion strategy, combining their complementary strengths to enhance the system’s overall robustness and effectiveness. Extensive evaluations on benchmark datasets, including UCSD Peds1, UCSD Peds2, CUHK Avenue, and ShanghaiTech, demonstrate that STAD-AI achieves superior performance with AUC scores of 86.6%, 99.1%, 91.4%, and 77.7%, respectively. These results highlight the framework’s ability to effectively leverage spatial and temporal features for detecting anomalies with high precision, advancing the state-of-the-art in video anomaly detection.