AudioCaps

dataset·audio captions·active

Crowdsourced natural-language captions paired with AudioSet-derived audio clips.

Recorded facts

Official sitehttps://audiocaps.github.io/
GeographySouth Korea
sizeaudio clips approx: 46000
scopeGeneral audio captioning including music and musical-event clips.
genresmixed audio; includes music
stewardAudioCaps authors
creatorsChris Dongjoo Kim; Byeongchang Kim; Hyunmin Lee; Gunhee Kim
languagesEnglish
modalitiesAudioSet clip references; human-written English captions; split metadata
geographiesGlobal
access methodOfficial data and code links.
intended usesaudio captioning; audio-language representation; music captioning transfer
related papersAudioCaps: Generating Captions for Audios in the Wild
annotation methodCaptions collected by crowdsourcing and checked for faithfulness.
collection methodAudio clips selected from AudioSet.
provenance claimsDerived from AudioSet, so audio access follows YouTube references.
current availabilityAnnotations available; source audio depends on AudioSet/YouTube.
us market scope basisPublicly available or materially used in US-facing music/audio-AI research.
limitations biases disputesGeneral-audio rather than music-specific.; Audio availability inherits AudioSet URL decay.
related models tools benchmarksAudioSet; WavCaps

Current

derived fromAudioSetSource 1 VERIFIED high confidence

Sources & changes

Checked 10d ago · highhow verification works
  • Status unknown → active
  • Official URL unknown → audiocaps.github.io
  • Record maintenance · 12 fields updated

Is this yours? Claim this record →·See something wrong? Report a correction →