Updated
Updated · arxiv.org · Sep 2
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Updated
Updated · arxiv.org · Sep 2

TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models

1 articles · Updated · arxiv.org · Sep 2

Summary

  • Researchers have introduced TAG-Bench, a benchmark for evaluating temporal audio grounding in large audio language models (LALMs).
  • TAG-Bench features 1,750 human-verified query–recording pairs across 149.5 hours of audio, testing when models can accurately localize queried content.
  • Results show current models struggle with precise localization and occurrence counting, highlighting significant challenges for practical audio search applications.