Sound Classification Using Whisper Embeddings

The ability for computers to recognize sounds from our environment provides valuable information for technology like robots/AI agents, smart devices, and surveillance cameras including helping robots assist the hard of hearing). As these systems become more autonomous, they need reliable methods to classify and distinguish thousands of different types of environmental sounds. This project develops a machine learning pipeline that is trained on embeddings of audio (potentially containing ambient noise, making speakers more difficult to resolve), clusters similar sounds, and classifies unseen speakers by their source.
Intern: Vivian Cai

Mentor: Francis Roxas (AMDS/A5E)