PhD defence Gijsbert Willem Adriaan Wijngaard

Supervisors: Prof. Dr. Elia Formisano, Prof. Dr. Michel Dumontier

Keywords: Audio-language models, Machine listening, Data curation, Semantic reasoning

 

"Data and Semantic Reasoning in Audio-language Learning"

 

This thesis investigates how computers can learn to understand everyday sounds and to describe them in ordinary language. Such systems learn from large collections of sound recordings paired with written descriptions. The research first maps 69 of these collections and shows that they often contain the same recordings, which distorts tests of how well the systems work. Cleaning up the examples and presenting easy ones before hard ones makes systems better at answering questions about sound. Inspired by how people describe what they hear, the thesis then presents a system that first reasons about who or what makes a sound, how it is produced, and where and when it happens, before it answers. It also introduces a method that judges computer-generated sound descriptions on meaning rather than exact wording, as people do. Finally, it shows that a coordinating program that consults multiple specialized listening programs, like a team of experts, produces the best results without extra training.

Click here for the live stream.