Industrial informatics

Autonomous Driving

The project investigates advanced visual-perception methods for autonomous vehicles and driver-assistance systems, where reliable interpretation of the surrounding environment is essential for safe navigation and timely decision-making. The research focuses on extracting geometric and semantic information directly from images acquired by conventional RGB cameras, with the aim of limiting the dependence on costly or complex sensing devices.

A first research direction addresses the detection of pedestrians and cyclists and the estimation of their distance from the vehicle. The proposed pipeline combines convolutional neural networks for semantic segmentation with monocular depth estimation, allowing vulnerable road users to be identified and their approximate distance to be evaluated from a single RGB image. This solution is designed to support driver-attention systems and collision-prevention functions without requiring LiDAR scanners or dedicated three-dimensional acquisition hardware. Experiments on public automotive datasets demonstrate the feasibility of jointly estimating object location and depth in complex road scenes.

A second direction concerns monocular depth estimation for reconstructing scene geometry from a single image. The project studies patch-based processing as an alternative to conventional full-image inference. In particular, a perspective-aware warp extraction strategy is introduced to compensate for camera distortions and to preserve fine spatial details at different positions in the image. The method can be integrated into both training and inference pipelines and is evaluated with multiple neural backbones and depth-estimation architectures. Results show consistent improvements over full-image processing and classical crop-based patch extraction across several error metrics, indicating that properly designed local processing can increase depth accuracy without requiring more expensive sensors.

The project also investigates explainable semantic segmentation for autonomous driving. A variational autoencoder architecture, Mgrad2VAE, generates both semantic segmentation masks and visual attention maps that reveal which image regions influence the model. Explainability is obtained through multiscale second-order derivatives between the latent representation and the encoder layers, capturing variations in the learned activations at different spatial scales.

Mgrad2VAE was evaluated on the SYNTHIA and A2D2 autonomous-driving datasets. The generated attention maps closely follow the corresponding semantic masks, while the model achieves AUC-ROC values of 83.20% on SYNTHIA and 95.36% on A2D2 for attention-based pixel classification. The proposed attention mechanism also improves segmentation performance compared with conventional deep variational autoencoders and an Xception-based model.

A related line of work develops Grad2VAE, an explainable variational autoencoder that uses second-order information to preserve the curvature of learned representations and generate online attention maps. Although applicable beyond autonomous driving, this research provides methodological foundations for interpreting unsupervised and generative models employed in complex visual environments.

Overall, the project combines semantic understanding, metric depth recovery, vulnerable-road-user analysis, and explainable artificial intelligence. Its main contribution is a set of image-based techniques that support accurate and interpretable environmental perception using affordable cameras, making them suitable for embedded, assisted-driving, and autonomous-navigation applications.

Project website

Relevant publications

A. Genovese, V. Piuri, F. Rundo, F. Scotti, C. Spampinato, Driver attention assistance by pedestrian/cyclist distance estimation from a single RGB Image: A CNN-based semantic segmentation approach, Proc. of the 22nd IEEE Int. Conf. on Industrial Technology (ICIT 2021), pp. 875-880, Valencia, Spain, March 2021, ISSN 978-1-7281-5730-6.
P. Coscia, A. Fusillo, A. Genovese, V. Piuri, F. Scotti, On the relevance of patch-based extraction methods for monocular depth estimation, Image and Vision Computing, vol. 166, no. 105857, pp. 1-17, February 2026, ISSN 0262-8856.
M. Abukmeil, S. Ferrari, A. Genovese, V. Piuri, F. Scotti, Grad2VAE: An explainable variational autoencoder model based on online attentions preserving curvatures of representations, Proc. of the 21st Int. Conf. on Image Analysis and Processing (ICIAP 2022), pp. 670–681, Lecce, Italy, May 2022, ISSN 978-3-031-06427-2.
M. Abukmeil, A. Genovese, V. Piuri, F. Rundo, F. Scotti, Towards explainable semantic segmentation for autonomous driving systems by multi-scale variational attention, Proc. of the 1st IEEE Int. Conf. on Autonomous Systems (ICAS 2021), pp. 1-5, Montreal, QC, Canada, August 2021, ISSN 978-1-7281-7289-7.