AudioCaps
dataset·audio captions·active
Crowdsourced natural-language captions paired with AudioSet-derived audio clips.
Recorded facts
| Official site | https://audiocaps.github.io/ ↗ |
|---|---|
| Geography | South Korea |
| size | audio clips approx: 46000 |
| scope | General audio captioning including music and musical-event clips. |
| genres | mixed audio; includes music |
| steward | AudioCaps authors |
| creators | Chris Dongjoo Kim; Byeongchang Kim; Hyunmin Lee; Gunhee Kim |
| languages | English |
| modalities | AudioSet clip references; human-written English captions; split metadata |
| geographies | Global |
| access method | Official data and code links. |
| intended uses | audio captioning; audio-language representation; music captioning transfer |
| related papers | AudioCaps: Generating Captions for Audios in the Wild |
| annotation method | Captions collected by crowdsourcing and checked for faithfulness. |
| collection method | Audio clips selected from AudioSet. |
| provenance claims | Derived from AudioSet, so audio access follows YouTube references. |
| current availability | Annotations available; source audio depends on AudioSet/YouTube. |
| us market scope basis | Publicly available or materially used in US-facing music/audio-AI research. |
| limitations biases disputes | General-audio rather than music-specific.; Audio availability inherits AudioSet URL decay. |
| related models tools benchmarks | AudioSet; WavCaps |
Current
| derived from | AudioSetSource 1 ↗ VERIFIED high confidence |
|---|
Sources & changes
Checked 10d ago · highhow verification works
Field-level evidence
Public change history
- Status unknown → active
- Official URL unknown → audiocaps.github.io
- Record maintenance · 12 fields updated
Is this yours? Claim this record →·See something wrong? Report a correction →
