Abstract
As collaborative robots (cobots) are increasingly integrated into unstructured environments, intelligent safety systems capable of detecting hazards without relying on extensive labeled data become necessary. This survey reviews deep learning methods for Industrial Visual Anomaly Detection (IVAD), emphasizing their role in robotic safety monitoring. Unlike static quality inspection, robotic IVAD must function in real time, handle dynamic scenes, and detect open-set anomalies. Existing approaches are categorized into reconstruction-based, embedding-based, and vision-language model (VLM) paradigms. Relevant benchmarks and robotics-specific datasets are examined, highlighting challenges such as motion, viewpoint shifts, and multimodal sensing. Strong benchmark scores do not carry over to robotic deployments, where dynamic scenes, latency limits, and semantic risks apply. This survey treats IVAD as a safety-critical system component, one that requires edge efficiency, low latency, and integration with middleware such as ROS2. The survey closes by mapping open problems and practical design directions for practitioners.