MMAU-Pro

dataset·general audio reasoning·active

Expert-created audio questions covering speech, sound, music, mixtures, spatial audio and long-form reasoning.

Recorded facts

Official sitehttps://sonalkum.github.io/mmau-pro/
GeographyUnited States
arxiv2508.13992
huggingface repositorygamma-lab-umd/MMAU-Pro
sizeskills: 49 · instances: 5305 · evaluated models: 22
scopeHolistic audio-general-intelligence evaluation with deliberately difficult questions.
genresmusic subset is multicultural but majority Western
licenseunknown/mixed upstream
stewardUniversity of Maryland-led Gamma Lab collaboration
creatorsSonal Kumar; Šimon Sedláček; Vaibhavi Lokegaonkar
languagesEnglish questions; multilingual audio included
modalitiessingle and multiple audio; multiple-choice QA; open-ended QA; spatial audio; long-form audio
geographiesGlobal
access methodOfficial page links dataset and code.
intended usesaudio reasoning evaluation; multimodal model evaluation
consent claimsSource-person and annotator consent details not verified.
related papersMMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
annotation methodHuman experts authored multi-hop questions and answers.
collection methodAudio was sourced from the wild rather than fixed benchmark corpora.
provenance claimsOfficial project describes wild-source collection and expert authorship.
current availabilityProject page and Hugging Face dataset available.
commercial use limitsWild-source media rights require separate assessment.
us market scope basisPublicly available or materially used in US-facing music/audio-AI research.
limitations biases disputesApproximately sixty percent of music content is Western.; Wild-source audio creates rights and stability risk.

Sources & changes

Checked 10d ago · highhow verification works
  • Status unknown → active
  • Official URL unknown → sonalkum.github.io/mmau-pro
  • Record maintenance · 12 fields updated

Is this yours? Claim this record →·See something wrong? Report a correction →