Datasets
102 records·Datasets, corpora and benchmarks for music and audio AI — with license, provenance and known limitations.
| Name | Type | Geography | Status | Checked |
|---|---|---|---|---|
| ACE-KiSing Scaled singing dataset introduced with ACE-Opencpop for multi-singer voice synthesis research. | large scale multilingual singing dataset | Beijing, Mainland China, China | active | Checked 10d ago |
| ACE-Opencpop Scaled multi-singer Mandarin singing dataset introduced for large-scale singing-voice synthesis. | large scale mandarin singing dataset | Beijing, Mainland China, China | active | Checked 10d ago |
| AI Music Tools Russian-language directory and practical resource for AI music, audio, DAW, analysis, conversion, rhyme, and prompt tools. | ai music resource directory | Russian-speaking market (operator location unverified) | active | Checked 10d ago |
| AIST Annotation for the RWC Music Database Manual beat, melody and chorus annotations for RWC musical pieces. | music structure annotation dataset | Tsukuba, Japan | active | Checked 10d ago |
| AcousticBrainz Crowdsourced low- and high-level acoustic descriptors indexed by MusicBrainz recording identifiers. | crowdsourced music features | Spain | active | Checked 10d ago |
| AcousticBrainz Genre Dataset Four genre-annotation datasets joining AcousticBrainz features with AllMusic, Discogs, Last.fm and Tagtraum labels. | genre benchmark | Spain | active | Checked 10d ago |
| Annotated Beethoven Piano Sonatas EPFL DCML corpus of annotated Beethoven piano-sonata scores for computational musicology. | annotated music score corpus | Lausanne, Romandy, Switzerland | active | Checked 10d ago |
| AudioCaps Crowdsourced natural-language captions paired with AudioSet-derived audio clips. | audio captions | South Korea | active | Checked 10d ago |
| AudioSet Human-labeled ten-second YouTube segments spanning music, speech and environmental sound classes. | audio event labels | United States | active | Checked 10d ago |
| Australian AI Music Alliance AI Music Charts Alliance-operated AI-music chart product; the accessible chart emphasized mainland-China tracks and exposes limited methodology. | directory or chart | Australia | unknown | Checked 10d ago |
| Azərbaycan Süni İntellekt Məzmunlarının Qorunması və Təhlili Mərkəzi Azerbaijani platform presenting registration, provenance, analysis, and protection services for AI-created or altered content including voice and music. | ai content rights registry and analysis center | Azerbaijan | unknown | Checked 10d ago |
| BAF: an audio fingerprinting dataset for broadcast monitoring UGent dataset pairing reference music with television query audio for broadcast fingerprinting research. | broadcast audio fingerprinting dataset | Belgium | active | Checked 10d ago |
| BAM! Artiestenmonitor Annual artist survey from BAM! Popauteurs and academic collaborators covering income, working conditions, and generative-AI use and concern. | annual artist monitor | Netherlands | active | Checked 10d ago |
| Bach Doodle Dataset User-entered melodies and model-generated four-voice harmonizations contributed through the 2019 Bach Doodle. | user contributed symbolic music | United States | active | Checked 10d ago |
| Brazilian Music Time-Series Dataset Conjunto de Dados de Séries Temporais de Músicas Brasileiras Dataset combining Spotify-derived acoustic features and processed lyrics for approximately 51,000 Brazilian tracks. | brazilian music features and lyrics dataset | Belo Horizonte, Brazil | active | Checked 10d ago |
| Brazilian Rhythmic Instruments Dataset Open MIR dataset of Brazilian rhythmic instruments, solo and multi-instrument recordings across five rhythm styles. | brazilian rhythmic instruments audio dataset | Rio de Janeiro, Brazil | active | Checked 10d ago |
| CMI-Bench A test-only benchmark converting diverse MIR tasks into standardized music instruction-following evaluation. | music instruction following | Global | active | Checked 10d ago |
| CMI-RewardBench Preference data and evaluation for reward models scoring musicality, text alignment and compositional instruction following. | music reward modeling | Global | active | Checked 10d ago |
| Cadenza Lyric Intelligibility Prediction Dataset Popular-music signals with lyrics and listener-derived intelligibility scores for the ICASSP 2026 Cadenza Challenge. | lyric intelligibility | United Kingdom | active | Checked 10d ago |
| CocoChorales Synthetic four-part chamber-ensemble music with mixtures, stems, MIDI and fine-grained performance annotations. | synthetic multitrack music | United States | active | Checked 10d ago |
| ComMU Pozalabs symbolic dataset of 11,144 professional-composer MIDI samples with 12 metadata fields. | symbolic music generation | South Korea | active | Checked 10d ago |
| CoversBR Large predominantly Brazilian music database for cover-song and live-song identification, with public metadata and features. | cover and live song identification dataset | Brazil | active | Checked 10d ago |
| CtrSVDD Controlled singing-voice deepfake corpus generated from public singing data with synthesis and voice-conversion systems. | controlled singing deepfake | Global | active | Checked 10d ago |
| DALI Full-song lyric, vocal-note and audio alignments at note, word, line and paragraph granularity. | aligned lyrics notes | France | active | Checked 10d ago |
| DCML Annotated Music Score Corpora Collection of harmonically and formally annotated symbolic music corpora. | annotated music score corpus collection | Lausanne, Switzerland | active | Checked 10d ago |
| DEAM Music excerpts and full songs with continuous and song-level valence and arousal annotations. | music emotion | Europe | active | Checked 10d ago |
| EMOPIA Pop-piano audio and MIDI clips labeled for perceived emotion, supporting recognition and emotion-conditioned generation. | music emotion audio midi | Taiwan | active | Checked 10d ago |
| EgoMusic Synchronized egocentric audiovisual dataset for music understanding and hearing-enhancement research. | egocentric audiovisual music dataset | Ireland | active | Checked 10d ago |
| Erkomaishvili Dataset Curated and annotated corpus of traditional Georgian vocal music for computational musicology and MIR tasks. | georgian vocal music corpus | Georgia / Georgian music | active | Checked 9d ago |
| Expanded Groove MIDI Dataset Expanded human drum-performance corpus with MIDI-aligned audio rendered across 43 drum kits. | aligned drum audio midi | United States | active | Checked 10d ago |
| FSD50K Open human-labeled Freesound clips across 200 AudioSet-ontology classes, including many musical sounds. | sound event audio | Spain | active | Checked 10d ago |
| FakeMusicCaps MusicCaps prompts regenerated with multiple text-to-music systems for synthetic-music detection and generator attribution. | synthetic music detection | Italy | active | Checked 10d ago |
| Free Music Archive Dataset Full-length Creative Commons music with audio, features, metadata and hierarchical genre labels for MIR. | creative commons music audio | Global | active | Checked 10d ago |
| GAPS Classical-guitar audio, scores and high-resolution MIDI alignments for transcription research. | guitar audio score alignment | United Kingdom | active | Checked 10d ago |
| GiantMIDI-Piano Machine-transcribed solo-piano MIDI derived from web recordings and IMSLP work metadata. | machine transcribed piano midi | Global | active | Checked 10d ago |
| Groove MIDI Dataset Human-performed expressive drum MIDI with aligned synthesized audio and performance metadata. | drum performance midi | United States | active | Checked 10d ago |
| GuitarSet Acoustic-guitar recordings with hexaphonic audio and time-aligned note, chord, beat and technique annotations. | guitar transcription | United States | active | Checked 10d ago |
| HEAR 2021 A NeurIPS challenge and evaluation suite for general-purpose audio representations across speech, sound and music tasks. | general audio representation | United States | active | Checked 10d ago |
| HF2 Hardanger Fiddle Dataset Paired audio and MIDI dataset of Norwegian Hardanger fiddle material for automatic music transcription research. | paired audio midi music dataset | Norway | active | Checked 10d ago |
| ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge A completed grand challenge for predicting overall and five fine-grained human aesthetic ratings of AI-generated songs. | song aesthetics prediction | Global | closed | Checked 10d ago |
| ICASSP 2026 Cadenza Challenge A completed challenge to predict lyric intelligibility for popular music, including listeners with hearing loss. | lyric intelligibility prediction | United Kingdom | closed | Checked 10d ago |
| IDMT-SMT-Chords Fraunhofer IDMT dataset for chord-recognition research. | chord recognition dataset | Germany | active | Checked 10d ago |
| IDMT-SMT-Drums Fraunhofer IDMT dataset for drum transcription research. | drum transcription dataset | Germany | active | Checked 10d ago |
| JVS-MuSiC Japanese multispeaker singing corpus with recordings from 100 singers. | japanese multispeaker singing voice corpus | Japan | active | Checked 10d ago |
| KazEmoTTS Kazakh emotional speech dataset for text-to-speech synthesis with six labeled emotions. | emotional tts dataset | Kazakhstan / Kazakh language | active | Checked 10d ago |
| Kazakh Songs ASR Dataset Gated non-commercial dataset of manually aligned Kazakh sung-vocal audio and text for ASR research. | sung speech asr dataset | Kazakhstan / Kazakh language | active | Checked 10d ago |
| Lakh MIDI Dataset Large deduplicated MIDI collection with matched and aligned subsets linked to the Million Song Dataset. | symbolic music | United States | active | Checked 10d ago |
| M6 Multi-generator, multi-domain, multilingual, multicultural, multi-genre and multi-instrument music-detection databases. | machine generated music detection | Global | active | Checked 10d ago |
| MADB Large music-aesthetics benchmark with multi-dimensional professional ratings and textual comments. | music aesthetics | Global | active | Checked 10d ago |
| MAESTRO Paired piano audio and high-precision MIDI from International Piano-e-Competition performances. | aligned piano audio midi | United States | active | Checked 10d ago |
| MARBLE A unified evaluation suite for pretrained music-audio representations across acoustic, performance, score and high-level tasks. | music audio representation | Global | active | Checked 10d ago |
| MID-FiLD Pozalabs MIDI dataset of 4,422 professional-writer samples for fine-level dynamics control. | expressive midi | South Korea | active | Checked 10d ago |
| MIR-1K Dataset of 1,000 Chinese karaoke song clips with separated accompaniment/vocal channels and manual annotations. | singing voice separation dataset | Taiwan | active | Checked 10d ago |
| MMAU A multi-domain benchmark of expert audio understanding and reasoning spanning speech, environmental sound and music. | general audio reasoning | United States | active | Checked 10d ago |
| MMAU-Pro Expert-created audio questions covering speech, sound, music, mixtures, spatial audio and long-form reasoning. | general audio reasoning | United States | active | Checked 10d ago |
| MTG-Jamendo Dataset | audio tags | unknown | active | Checked 11d ago |
| MUSDB18 Full-length stereo music mixtures with isolated drums, bass, vocals and other stems for source separation. | music source separation | Global | active | Checked 10d ago |
| MUSIB Open benchmark and software package for reproducible evaluation of musical-score inpainting systems. | musical score inpainting benchmark | Chile | active | Checked 10d ago |
| MagnaTagATune Magnatune audio clips with human tags collected through the TagATune game for music auto-tagging research. | music autotagging | Global | active | Checked 10d ago |
| McGill Billboard Project Chord, structure and feature annotations for a sampled set of Billboard-chart songs. | chord structure annotations | Canada | active | Checked 10d ago |
| MedleyDB Annotated royalty-free multitrack recordings supporting melody, pitch, instrument and source-separation research. | multitrack music | United States | active | Checked 10d ago |
| Million Song Dataset Audio features and metadata for one million contemporary popular-music tracks, without distributed audio. | music features metadata | United States | active | Checked 10d ago |
| MoisesDB Multitrack music dataset with a hierarchical stem taxonomy for fine-grained source-separation research. | multitrack source separation dataset | Brazil | active | Checked 10d ago |
| MuChoMusic Human-validated multiple-choice questions for evaluating music understanding and reasoning in audio-language models. | music audio language qa | Global | active | Checked 10d ago |
| Music Arena A live pairwise evaluation platform and rolling leaderboard for text-to-music generation systems. | live text to music preference | United States | active | Checked 10d ago |
| Music Arena Dataset Rolling human preference data from live pairwise comparisons of text-to-music model outputs. | human music preferences | United States | active | Checked 10d ago |
| Music4All-Onion Multimodal music features and large-scale Last.fm listening histories extending Music4All. | multimodal music recommendation | Europe | active | Checked 10d ago |
| MusicBench Music audio-text pairs expanding MusicCaps with extracted musical features and augmentation for controllable generation. | text to music training | Singapore | active | Checked 10d ago |
| MusicCaps Dataset of 5.5k music clips with expert-written text captions | audio captions | unknown | active | Checked 11d ago |
| MusicNet Classical recordings aligned to score-derived note, instrument and metrical-position labels. | classical note annotation | United States | active | Checked 10d ago |
| MusicTGA-HR Rights-cleared music data infrastructure for AI services, music supply and API integration. | rights cleared music data and api infrastructure | Tokyo, Japan | active | Checked 10d ago |
| NSynth Dataset Annotated four-second musical-note audio for timbre modeling, synthesis, and audio representation research. | instrument note audio | United States | active | Checked 10d ago |
| NVPI Muziekmonitor Annual Dutch music-consumer monitor commissioned or published by NVPI, with recurring measures of attitudes toward AI-generated music. | annual music consumer monitor | Netherlands | active | Checked 10d ago |
| OpenMIC-2018 Creative Commons music excerpts partially labeled for presence or absence of 20 instrument classes. | instrument recognition | United States | active | Checked 10d ago |
| Opencpop Studio-recorded Mandarin singing corpus with 100 songs and phoneme, note and pitch annotations. | mandarin singing voice corpus | Xi'an, Mainland China, China | active | Checked 10d ago |
| OrchideaSOL Subset of the IRCAM Studio Online instrument-note collection prepared for orchestration research. | instrument note recording dataset | Paris, France | active | Checked 10d ago |
| PAN-AR Dataset of higher-order ambisonics room impulse responses, ambient noise and spherical pictures. | higher order ambisonics multimodal dataset | Milan, Lombardy, Italy | active | Checked 10d ago |
| PDMX Large MusicXML collection of scores labeled public domain, with score and user-interaction metadata. | public domain musicxml | United States | active | Checked 10d ago |
| PJS Phoneme-balanced Japanese singing-voice corpus licensed CC BY-SA 4.0. | japanese singing voice corpus | Japan | active | Checked 10d ago |
| POP909 Professional piano arrangements of 909 popular songs with aligned MIDI and structural annotations. | symbolic pop arrangement | China | active | Checked 10d ago |
| RWC Music Database AIST research database containing six collections and 315 recorded musical pieces. | copyright cleared music research database | Tsukuba, Japan | active | Checked 10d ago |
| SALAMI Hierarchical expert structural annotations for a large and stylistically diverse music collection. | music structure annotations | Canada | active | Checked 10d ago |
| SAMBASET Dataset of historical samba-enredo recordings created for MIR research on Brazilian music. | historical samba enredo dataset | Rio de Janeiro, Brazil | active | Checked 10d ago |
| SVDD Challenge 2024 The inaugural controlled and in-the-wild singing-voice deepfake detection challenge held with IEEE SLT 2024. | singing voice deepfake detection | Global | closed | Checked 10d ago |
| Saraga collections Open annotated audio corpora for Carnatic and Hindustani music research. | annotated music corpus | India | active | Checked 10d ago |
| SingFake In-the-wild real and deepfake singing data for singing-voice authenticity detection. | singing voice deepfake | United States | active | Checked 10d ago |
| Slakh2100 Synthetic multitrack audio and aligned MIDI rendered from Lakh MIDI files for separation and transcription. | synthetic multitrack music | United States | active | Checked 10d ago |
| Song Describer Dataset Crowdsourced human descriptions of Creative Commons music for music-language evaluation. | music captions | Global | active | Checked 10d ago |
| SongCompose-PT Chinese-English pretraining dataset containing lyrics, melodies and paired lyric-melody data for SongComposer. | bilingual lyrics melody pretraining dataset | Mainland China, China | active | Checked 10d ago |
| SongEval Full-length real and generated songs with expert ratings across five aesthetic dimensions. | song aesthetics | China | active | Checked 10d ago |
| Sound Demixing Challenge 2023 A completed AIcrowd challenge with music and cinematic source-separation tracks and hidden evaluation sets. | sound demixing | Global | closed | Checked 10d ago |
| Sounds Queer Dataset of AI-generated music from queer-identity prompts for auditing text-to-music models. | text to music audit dataset | Germany | active | Checked 10d ago |
| Spotify Million Playlist Dataset One million anonymized public playlists released for automatic playlist continuation research. | playlist recommendation | United States | active | Checked 10d ago |
| Suno Music Generation Dataset Metadata and downloadable links for 659,788 Suno-generated songs discovered through systematic search queries. | ai generated music links | Global | active | Checked 10d ago |
| Teach Yourself Georgian Folk Songs Dataset Annotated corpus of traditional Georgian vocal polyphony intended for computational musicology and MIR research. | annotated traditional vocal polyphony | Georgia / Georgian music | active | Checked 9d ago |
| TinySOL Compact dataset of isolated instrumental notes published through IRCAM Forum. | instrument note recording dataset | Paris, France | active | Checked 10d ago |
| Tohoku Kiritan Singing Database 東北きりたん歌唱データベース Research singing database used to create the Tohoku Kiritan library distributed with NEUTRINO. | japanese character singing database | Japan | unknown | Checked 10d ago |
| Tunepal tune corpus Search corpus of more than 23,000 Irish traditional music tunes used by Tunepal. | traditional music tune corpus | Ireland | active | Checked 10d ago |
| URMP Dataset Coordinated multi-instrument classical performances with separated audio, assembled video, scores and pitch annotations. | multimodal music performance | United States | active | Checked 10d ago |
| VocalSet A cappella singing recordings across singers, vowels, registers and vocal techniques, with a corrected annotated derivative. | singing voice | United States | active | Checked 10d ago |
| WavCaps Large weakly labeled audio-caption corpus assembled from web sound libraries and AudioSet, with ChatGPT-assisted caption cleaning. | weak audio captions | United Kingdom | active | Checked 10d ago |
| WildSVDD In-the-wild singing voice deepfake detection data extending SingFake for the SVDD 2024 challenge. | in the wild singing deepfake | United States | active | Checked 10d ago |
