I am a member of technical staff at AMI Labs building world models. Previously, I was a Staff Research Scientist at Meta Superintelligence Labs and Facebook AI Research (FAIR), where I was a core contributor to Meta's speech and audio foundation models including SAM Audio, MovieGen Audio, AudioBox, VoiceBox, and MMS. I obtained my Ph.D. from TTIC where I worked on automatic sign language understanding under the advisement of Prof. Karen Livescu.
June 2026 — After 4 amazing years, I moved from Meta to Advanced Machine Intelligence Labs.
December 2025 — Launched SAM Audio, a foundation model that extends Segment Anything to audio, enabling general-purpose audio separation via multimodal prompts.
March 2025 — Our team released AudioBox-aesthetics, a unified automatic quality assessment framework for any audio.