FoleyBench: A Benchmark for Video-to-Audio Models
The first large-scale benchmark for Foley-style video-to-audio generation: non-speech, non-music sound that must align causally with the visuals.
I'm a PhD student at Carnegie Mellon University, advised by Chris Donahue and Bhiksha Raj.
My research is on multimodal intelligence, particularly audio, and on reliable ways to evaluate it.
I've worked on:
Previously, I was a founding engineer at Cekura (YC F24), where I built the evaluation layer for voice agents, which now runs on 60K+ calls a day. Before that, I completed an MS at CMU, research internships at MIT and EPFL, and a BTech in electrical engineering from IIT Delhi.
The first large-scale benchmark for Foley-style video-to-audio generation: non-speech, non-music sound that must align causally with the visuals.
A benchmark for audio general intelligence across speech, sound and music, including spatial, multi-audio, and long-form questions that need several reasoning steps at once.
A metric for open-ended audio question answering that pairs LLM reasoning with an audio-entailment check. State-of-the-art correlation with human judgments on AQEval, a new 10k-response benchmark.
A small audio-language model that competes with much larger models on audio reasoning tasks, together with ReasonAQA, a large synthetic dataset for audio reasoning.
Evaluates whether vision-language models such as GPT-4o and Gemini can classify sounds from spectrogram images alone. Few-shot, they reach expert-level accuracy.