TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Updated
Updated · arxiv.org · Sep 2
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
1 articles · Updated · arxiv.org · Sep 2
Summary
Researchers have introduced TAG-Bench, a benchmark for evaluating temporal audio grounding in large audio language models (LALMs).
TAG-Bench features 1,750 human-verified query–recording pairs across 149.5 hours of audio, testing when models can accurately localize queried content.
Results show current models struggle with precise localization and occurrence counting, highlighting significant challenges for practical audio search applications.