PhD defence Alexandros Doumanoglou

Supervisors: Prof. Dr. Gerhard Weiss, Dr. Ir. Kurt Driessens, Dr. Stelios Asteriadis

Co-supervisor: Dr. Dimitris Zarpalas

Keywords: Neural Networks, Explainable AI, Mechanistic Interpretability, Computer Vision

 

"Learning Latent Space Directions for Mechanistic Interpretability of Deep Neural Networks"

 

Significant progress in Deep Learning has made Artificial Intelligence an integral part of daily life, offering solutions to challenges that were once deemed insurmountable to solve using standard software development practices. Specific examples include image recognition, object localization and semantic segmentation with applications in robotic surgeries, car plate recognition, self-driving cars and surveillance. In contrast to traditional computer programs that are comprised of a known instruction set, the instructions and the program memory of neural programs implemented by neural networks are hardly understood by humans. This thesis tries to shed light on how neural networks store and retrieve semantic information from their program’s memory, uncovering both the instructions and the memory locations where semantic variables are stored. This allows AI developers to debug neural programs, understand failure cases, and correct errors in the instructions, ensuring reliability, safety and improved user experience in AI systems. To add more, access to the neural program’s instructions and memory layout allows to obtain explanations regarding a neural network’s decisions, a critical requirement when those AI systems assist in healthcare, law enforcement and high-stakes decisions. This thesis not only investigates how neural programs read and write semantic information in their memory but takes a leap forward to demonstrate their utility in practical applications, adding one more tiny bit of knowledge in mechanistic interpretability research of neural networks.

Click here for the live stream.