Enrollment-free target speech extraction
Pulling out the person you're engaging with — without asking them for a reference clip first — so the system stays out of the way of the conversation.
Teaching machines to hear the one voice you're talking to.
Speech & audio AI + HCI. I build target speech extraction, speaker representations and real-time enhancement models that work in the noisy rooms where conversations actually happen. Advised by Dhruv Jain, with Hao-Wen Dong.
↳ move over the room: the voice nearest your cursor is extracted
Robust speech and audio perception in real-world conditions, from the model to the person wearing it.
Pulling out the person you're engaging with — without asking them for a reference clip first — so the system stays out of the way of the conversation.
Learning persistent speaker identities directly from overlapped, multi-talker audio instead of clean single-speaker recordings.
Low-latency models on raw waveforms, and systems that let people turn down the sounds that overwhelm them.
FNU Sidharth, Meysam Asgari, Hao-Wen Dong, Dhruv Jain
Jeremy Zhengqi Huang, Emani Hicks, Sidharth, Gillian R. Hayes, Dhruv Jain
Yan Ru Pei, Ritik Shrivastava, FNU Sidharth
Sidharth, Vishwas Sathish, Shweta Bansal, Samantha Sun, Timmy Pham, Kurt Weaver, Rajesh P. N. Rao, Jeffrey Herron
Shikha Baghel, Shreyas Ramoji, Sidharth, et al.
Sidharth, et al.
Jerrin Thomas Panachakel, H. Ranjana, Sidharth, et al.
K. Sana Parveen, Jerrin Thomas Panachakel, H. Ranjana, Sidharth, et al.
Submitted Unmixing the Crowd to IEEE SLT 2026.
Joined Bose Research for the summer — open-vocabulary, queryable sound event localization.
Sona appears at CHI 2026 Extended Abstracts.
Started the PhD in CSE at the University of Michigan.
aTENNuate accepted at Interspeech 2025.
Presented PainDECOG at AAAI 2025 W3PHIAI.
Decoding Pain: Statistical Identification of Biomarkers from Electrophysiological SignalsAAAI 2025 · W3PHIAI workshop · Mar 2025
Emotion Detection from EEG using Transfer LearningIEEE EMBC 2023 · Sydney · Jul 2023