Datasets

102 records·Datasets, corpora and benchmarks for music and audio AI — with license, provenance and known limitations.

NameTypeGeographyStatusChecked
ACE-KiSing
Scaled singing dataset introduced with ACE-Opencpop for multi-singer voice synthesis research.
large scale multilingual singing datasetBeijing, Mainland China, ChinaactiveChecked 10d ago
ACE-Opencpop
Scaled multi-singer Mandarin singing dataset introduced for large-scale singing-voice synthesis.
large scale mandarin singing datasetBeijing, Mainland China, ChinaactiveChecked 10d ago
AI Music Tools
Russian-language directory and practical resource for AI music, audio, DAW, analysis, conversion, rhyme, and prompt tools.
ai music resource directoryRussian-speaking market (operator location unverified)activeChecked 10d ago
AIST Annotation for the RWC Music Database
Manual beat, melody and chorus annotations for RWC musical pieces.
music structure annotation datasetTsukuba, JapanactiveChecked 10d ago
AcousticBrainz
Crowdsourced low- and high-level acoustic descriptors indexed by MusicBrainz recording identifiers.
crowdsourced music featuresSpainactiveChecked 10d ago
AcousticBrainz Genre Dataset
Four genre-annotation datasets joining AcousticBrainz features with AllMusic, Discogs, Last.fm and Tagtraum labels.
genre benchmarkSpainactiveChecked 10d ago
Annotated Beethoven Piano Sonatas
EPFL DCML corpus of annotated Beethoven piano-sonata scores for computational musicology.
annotated music score corpusLausanne, Romandy, SwitzerlandactiveChecked 10d ago
AudioCaps
Crowdsourced natural-language captions paired with AudioSet-derived audio clips.
audio captionsSouth KoreaactiveChecked 10d ago
AudioSet
Human-labeled ten-second YouTube segments spanning music, speech and environmental sound classes.
audio event labelsUnited StatesactiveChecked 10d ago
Australian AI Music Alliance AI Music Charts
Alliance-operated AI-music chart product; the accessible chart emphasized mainland-China tracks and exposes limited methodology.
directory or chartAustraliaunknownChecked 10d ago
Azərbaycan Süni İntellekt Məzmunlarının Qorunması və Təhlili Mərkəzi
Azerbaijani platform presenting registration, provenance, analysis, and protection services for AI-created or altered content including voice and music.
ai content rights registry and analysis centerAzerbaijanunknownChecked 10d ago
BAF: an audio fingerprinting dataset for broadcast monitoring
UGent dataset pairing reference music with television query audio for broadcast fingerprinting research.
broadcast audio fingerprinting datasetBelgiumactiveChecked 10d ago
BAM! Artiestenmonitor
Annual artist survey from BAM! Popauteurs and academic collaborators covering income, working conditions, and generative-AI use and concern.
annual artist monitorNetherlandsactiveChecked 10d ago
Bach Doodle Dataset
User-entered melodies and model-generated four-voice harmonizations contributed through the 2019 Bach Doodle.
user contributed symbolic musicUnited StatesactiveChecked 10d ago
Brazilian Music Time-Series Dataset
Conjunto de Dados de Séries Temporais de Músicas Brasileiras
Dataset combining Spotify-derived acoustic features and processed lyrics for approximately 51,000 Brazilian tracks.
brazilian music features and lyrics datasetBelo Horizonte, BrazilactiveChecked 10d ago
Brazilian Rhythmic Instruments Dataset
Open MIR dataset of Brazilian rhythmic instruments, solo and multi-instrument recordings across five rhythm styles.
brazilian rhythmic instruments audio datasetRio de Janeiro, BrazilactiveChecked 10d ago
CMI-Bench
A test-only benchmark converting diverse MIR tasks into standardized music instruction-following evaluation.
music instruction followingGlobalactiveChecked 10d ago
CMI-RewardBench
Preference data and evaluation for reward models scoring musicality, text alignment and compositional instruction following.
music reward modelingGlobalactiveChecked 10d ago
Cadenza Lyric Intelligibility Prediction Dataset
Popular-music signals with lyrics and listener-derived intelligibility scores for the ICASSP 2026 Cadenza Challenge.
lyric intelligibilityUnited KingdomactiveChecked 10d ago
CocoChorales
Synthetic four-part chamber-ensemble music with mixtures, stems, MIDI and fine-grained performance annotations.
synthetic multitrack musicUnited StatesactiveChecked 10d ago
ComMU
Pozalabs symbolic dataset of 11,144 professional-composer MIDI samples with 12 metadata fields.
symbolic music generationSouth KoreaactiveChecked 10d ago
CoversBR
Large predominantly Brazilian music database for cover-song and live-song identification, with public metadata and features.
cover and live song identification datasetBrazilactiveChecked 10d ago
CtrSVDD
Controlled singing-voice deepfake corpus generated from public singing data with synthesis and voice-conversion systems.
controlled singing deepfakeGlobalactiveChecked 10d ago
DALI
Full-song lyric, vocal-note and audio alignments at note, word, line and paragraph granularity.
aligned lyrics notesFranceactiveChecked 10d ago
DCML Annotated Music Score Corpora
Collection of harmonically and formally annotated symbolic music corpora.
annotated music score corpus collectionLausanne, SwitzerlandactiveChecked 10d ago
DEAM
Music excerpts and full songs with continuous and song-level valence and arousal annotations.
music emotionEuropeactiveChecked 10d ago
EMOPIA
Pop-piano audio and MIDI clips labeled for perceived emotion, supporting recognition and emotion-conditioned generation.
music emotion audio midiTaiwanactiveChecked 10d ago
EgoMusic
Synchronized egocentric audiovisual dataset for music understanding and hearing-enhancement research.
egocentric audiovisual music datasetIrelandactiveChecked 10d ago
Erkomaishvili Dataset
Curated and annotated corpus of traditional Georgian vocal music for computational musicology and MIR tasks.
georgian vocal music corpusGeorgia / Georgian musicactiveChecked 9d ago
Expanded Groove MIDI Dataset
Expanded human drum-performance corpus with MIDI-aligned audio rendered across 43 drum kits.
aligned drum audio midiUnited StatesactiveChecked 10d ago
FSD50K
Open human-labeled Freesound clips across 200 AudioSet-ontology classes, including many musical sounds.
sound event audioSpainactiveChecked 10d ago
FakeMusicCaps
MusicCaps prompts regenerated with multiple text-to-music systems for synthetic-music detection and generator attribution.
synthetic music detectionItalyactiveChecked 10d ago
Free Music Archive Dataset
Full-length Creative Commons music with audio, features, metadata and hierarchical genre labels for MIR.
creative commons music audioGlobalactiveChecked 10d ago
GAPS
Classical-guitar audio, scores and high-resolution MIDI alignments for transcription research.
guitar audio score alignmentUnited KingdomactiveChecked 10d ago
GiantMIDI-Piano
Machine-transcribed solo-piano MIDI derived from web recordings and IMSLP work metadata.
machine transcribed piano midiGlobalactiveChecked 10d ago
Groove MIDI Dataset
Human-performed expressive drum MIDI with aligned synthesized audio and performance metadata.
drum performance midiUnited StatesactiveChecked 10d ago
GuitarSet
Acoustic-guitar recordings with hexaphonic audio and time-aligned note, chord, beat and technique annotations.
guitar transcriptionUnited StatesactiveChecked 10d ago
HEAR 2021
A NeurIPS challenge and evaluation suite for general-purpose audio representations across speech, sound and music tasks.
general audio representationUnited StatesactiveChecked 10d ago
HF2 Hardanger Fiddle Dataset
Paired audio and MIDI dataset of Norwegian Hardanger fiddle material for automatic music transcription research.
paired audio midi music datasetNorwayactiveChecked 10d ago
ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
A completed grand challenge for predicting overall and five fine-grained human aesthetic ratings of AI-generated songs.
song aesthetics predictionGlobalclosedChecked 10d ago
ICASSP 2026 Cadenza Challenge
A completed challenge to predict lyric intelligibility for popular music, including listeners with hearing loss.
lyric intelligibility predictionUnited KingdomclosedChecked 10d ago
IDMT-SMT-Chords
Fraunhofer IDMT dataset for chord-recognition research.
chord recognition datasetGermanyactiveChecked 10d ago
IDMT-SMT-Drums
Fraunhofer IDMT dataset for drum transcription research.
drum transcription datasetGermanyactiveChecked 10d ago
JVS-MuSiC
Japanese multispeaker singing corpus with recordings from 100 singers.
japanese multispeaker singing voice corpusJapanactiveChecked 10d ago
KazEmoTTS
Kazakh emotional speech dataset for text-to-speech synthesis with six labeled emotions.
emotional tts datasetKazakhstan / Kazakh languageactiveChecked 10d ago
Kazakh Songs ASR Dataset
Gated non-commercial dataset of manually aligned Kazakh sung-vocal audio and text for ASR research.
sung speech asr datasetKazakhstan / Kazakh languageactiveChecked 10d ago
Lakh MIDI Dataset
Large deduplicated MIDI collection with matched and aligned subsets linked to the Million Song Dataset.
symbolic musicUnited StatesactiveChecked 10d ago
M6
Multi-generator, multi-domain, multilingual, multicultural, multi-genre and multi-instrument music-detection databases.
machine generated music detectionGlobalactiveChecked 10d ago
MADB
Large music-aesthetics benchmark with multi-dimensional professional ratings and textual comments.
music aestheticsGlobalactiveChecked 10d ago
MAESTRO
Paired piano audio and high-precision MIDI from International Piano-e-Competition performances.
aligned piano audio midiUnited StatesactiveChecked 10d ago
MARBLE
A unified evaluation suite for pretrained music-audio representations across acoustic, performance, score and high-level tasks.
music audio representationGlobalactiveChecked 10d ago
MID-FiLD
Pozalabs MIDI dataset of 4,422 professional-writer samples for fine-level dynamics control.
expressive midiSouth KoreaactiveChecked 10d ago
MIR-1K
Dataset of 1,000 Chinese karaoke song clips with separated accompaniment/vocal channels and manual annotations.
singing voice separation datasetTaiwanactiveChecked 10d ago
MMAU
A multi-domain benchmark of expert audio understanding and reasoning spanning speech, environmental sound and music.
general audio reasoningUnited StatesactiveChecked 10d ago
MMAU-Pro
Expert-created audio questions covering speech, sound, music, mixtures, spatial audio and long-form reasoning.
general audio reasoningUnited StatesactiveChecked 10d ago
MTG-Jamendo Datasetaudio tagsunknownactiveChecked 11d ago
MUSDB18
Full-length stereo music mixtures with isolated drums, bass, vocals and other stems for source separation.
music source separationGlobalactiveChecked 10d ago
MUSIB
Open benchmark and software package for reproducible evaluation of musical-score inpainting systems.
musical score inpainting benchmarkChileactiveChecked 10d ago
MagnaTagATune
Magnatune audio clips with human tags collected through the TagATune game for music auto-tagging research.
music autotaggingGlobalactiveChecked 10d ago
McGill Billboard Project
Chord, structure and feature annotations for a sampled set of Billboard-chart songs.
chord structure annotationsCanadaactiveChecked 10d ago
MedleyDB
Annotated royalty-free multitrack recordings supporting melody, pitch, instrument and source-separation research.
multitrack musicUnited StatesactiveChecked 10d ago
Million Song Dataset
Audio features and metadata for one million contemporary popular-music tracks, without distributed audio.
music features metadataUnited StatesactiveChecked 10d ago
MoisesDB
Multitrack music dataset with a hierarchical stem taxonomy for fine-grained source-separation research.
multitrack source separation datasetBrazilactiveChecked 10d ago
MuChoMusic
Human-validated multiple-choice questions for evaluating music understanding and reasoning in audio-language models.
music audio language qaGlobalactiveChecked 10d ago
Music Arena
A live pairwise evaluation platform and rolling leaderboard for text-to-music generation systems.
live text to music preferenceUnited StatesactiveChecked 10d ago
Music Arena Dataset
Rolling human preference data from live pairwise comparisons of text-to-music model outputs.
human music preferencesUnited StatesactiveChecked 10d ago
Music4All-Onion
Multimodal music features and large-scale Last.fm listening histories extending Music4All.
multimodal music recommendationEuropeactiveChecked 10d ago
MusicBench
Music audio-text pairs expanding MusicCaps with extracted musical features and augmentation for controllable generation.
text to music trainingSingaporeactiveChecked 10d ago
MusicCaps
Dataset of 5.5k music clips with expert-written text captions
audio captionsunknownactiveChecked 11d ago
MusicNet
Classical recordings aligned to score-derived note, instrument and metrical-position labels.
classical note annotationUnited StatesactiveChecked 10d ago
MusicTGA-HR
Rights-cleared music data infrastructure for AI services, music supply and API integration.
rights cleared music data and api infrastructureTokyo, JapanactiveChecked 10d ago
NSynth Dataset
Annotated four-second musical-note audio for timbre modeling, synthesis, and audio representation research.
instrument note audioUnited StatesactiveChecked 10d ago
NVPI Muziekmonitor
Annual Dutch music-consumer monitor commissioned or published by NVPI, with recurring measures of attitudes toward AI-generated music.
annual music consumer monitorNetherlandsactiveChecked 10d ago
OpenMIC-2018
Creative Commons music excerpts partially labeled for presence or absence of 20 instrument classes.
instrument recognitionUnited StatesactiveChecked 10d ago
Opencpop
Studio-recorded Mandarin singing corpus with 100 songs and phoneme, note and pitch annotations.
mandarin singing voice corpusXi'an, Mainland China, ChinaactiveChecked 10d ago
OrchideaSOL
Subset of the IRCAM Studio Online instrument-note collection prepared for orchestration research.
instrument note recording datasetParis, FranceactiveChecked 10d ago
PAN-AR
Dataset of higher-order ambisonics room impulse responses, ambient noise and spherical pictures.
higher order ambisonics multimodal datasetMilan, Lombardy, ItalyactiveChecked 10d ago
PDMX
Large MusicXML collection of scores labeled public domain, with score and user-interaction metadata.
public domain musicxmlUnited StatesactiveChecked 10d ago
PJS
Phoneme-balanced Japanese singing-voice corpus licensed CC BY-SA 4.0.
japanese singing voice corpusJapanactiveChecked 10d ago
POP909
Professional piano arrangements of 909 popular songs with aligned MIDI and structural annotations.
symbolic pop arrangementChinaactiveChecked 10d ago
RWC Music Database
AIST research database containing six collections and 315 recorded musical pieces.
copyright cleared music research databaseTsukuba, JapanactiveChecked 10d ago
SALAMI
Hierarchical expert structural annotations for a large and stylistically diverse music collection.
music structure annotationsCanadaactiveChecked 10d ago
SAMBASET
Dataset of historical samba-enredo recordings created for MIR research on Brazilian music.
historical samba enredo datasetRio de Janeiro, BrazilactiveChecked 10d ago
SVDD Challenge 2024
The inaugural controlled and in-the-wild singing-voice deepfake detection challenge held with IEEE SLT 2024.
singing voice deepfake detectionGlobalclosedChecked 10d ago
Saraga collections
Open annotated audio corpora for Carnatic and Hindustani music research.
annotated music corpusIndiaactiveChecked 10d ago
SingFake
In-the-wild real and deepfake singing data for singing-voice authenticity detection.
singing voice deepfakeUnited StatesactiveChecked 10d ago
Slakh2100
Synthetic multitrack audio and aligned MIDI rendered from Lakh MIDI files for separation and transcription.
synthetic multitrack musicUnited StatesactiveChecked 10d ago
Song Describer Dataset
Crowdsourced human descriptions of Creative Commons music for music-language evaluation.
music captionsGlobalactiveChecked 10d ago
SongCompose-PT
Chinese-English pretraining dataset containing lyrics, melodies and paired lyric-melody data for SongComposer.
bilingual lyrics melody pretraining datasetMainland China, ChinaactiveChecked 10d ago
SongEval
Full-length real and generated songs with expert ratings across five aesthetic dimensions.
song aestheticsChinaactiveChecked 10d ago
Sound Demixing Challenge 2023
A completed AIcrowd challenge with music and cinematic source-separation tracks and hidden evaluation sets.
sound demixingGlobalclosedChecked 10d ago
Sounds Queer
Dataset of AI-generated music from queer-identity prompts for auditing text-to-music models.
text to music audit datasetGermanyactiveChecked 10d ago
Spotify Million Playlist Dataset
One million anonymized public playlists released for automatic playlist continuation research.
playlist recommendationUnited StatesactiveChecked 10d ago
Suno Music Generation Dataset
Metadata and downloadable links for 659,788 Suno-generated songs discovered through systematic search queries.
ai generated music linksGlobalactiveChecked 10d ago
Teach Yourself Georgian Folk Songs Dataset
Annotated corpus of traditional Georgian vocal polyphony intended for computational musicology and MIR research.
annotated traditional vocal polyphonyGeorgia / Georgian musicactiveChecked 9d ago
TinySOL
Compact dataset of isolated instrumental notes published through IRCAM Forum.
instrument note recording datasetParis, FranceactiveChecked 10d ago
Tohoku Kiritan Singing Database
東北きりたん歌唱データベース
Research singing database used to create the Tohoku Kiritan library distributed with NEUTRINO.
japanese character singing databaseJapanunknownChecked 10d ago
Tunepal tune corpus
Search corpus of more than 23,000 Irish traditional music tunes used by Tunepal.
traditional music tune corpusIrelandactiveChecked 10d ago
URMP Dataset
Coordinated multi-instrument classical performances with separated audio, assembled video, scores and pitch annotations.
multimodal music performanceUnited StatesactiveChecked 10d ago
VocalSet
A cappella singing recordings across singers, vowels, registers and vocal techniques, with a corrected annotated derivative.
singing voiceUnited StatesactiveChecked 10d ago
WavCaps
Large weakly labeled audio-caption corpus assembled from web sound libraries and AudioSet, with ChatGPT-assisted caption cleaning.
weak audio captionsUnited KingdomactiveChecked 10d ago
WildSVDD
In-the-wild singing voice deepfake detection data extending SingFake for the SVDD 2024 challenge.
in the wild singing deepfakeUnited StatesactiveChecked 10d ago