Models
110 records·Named models and versions behind the tools — capabilities, access, training and licensing claims as documented.
| Name | Type | Geography | Status | Checked |
|---|---|---|---|---|
| ACE-Step 1.5 Base Diffusion Transformer base checkpoint | generative music model | unknown | active | Checked 11d ago |
| ACE-Step 1.5 LM 0.6B Music-planning language model | music planning model | unknown | active | Checked 11d ago |
| ACE-Step 1.5 LM 1.7B Mid-size music-planning language model | music planning model | unknown | active | Checked 11d ago |
| ACE-Step 1.5 LM 4B Largest released music-planning language model | music planning model | unknown | active | Checked 11d ago |
| ACE-Step 1.5 SFT Supervised-fine-tuned checkpoint | generative music model | unknown | active | Checked 11d ago |
| ACE-Step 1.5 Turbo Eight-step accelerated checkpoint | generative music model | unknown | active | Checked 11d ago |
| ACRCloud Identification API v1 Signed multipart API for identifying audio or locally extracted fingerprints against recognition projects. | music recognition api | unknown | active | Checked 11d ago |
| Adobe Firefly Audio Model — Generate Soundtrack beta Adobe's beta soundtrack-generation model for guided background music and video matching. | generative music model | unknown | active | Checked 11d ago |
| Audio-MAGNeT Medium Non-autoregressive MAGNeT checkpoint for 10 seconds of sound effects. | generative audio model | unknown | active | Checked 11d ago |
| Audio-MAGNeT Small Non-autoregressive MAGNeT checkpoint for 10 seconds of sound effects. | generative audio model | unknown | active | Checked 11d ago |
| AudioGen Medium v2 Meta's released 1.5B text-to-sound checkpoint, identified as AudioGen model version 2. | generative audio model | unknown | active | Checked 11d ago |
| AudioLDM Latent-diffusion model for text-conditioned audio generation developed with University of Surrey participation. | text to audio model | United Kingdom research participation | active | Checked 10d ago |
| AudioLDM 2 Self-supervised foundation model for text-conditioned generation of speech, sound and music. | general audio generation model | United Kingdom research participation | active | Checked 10d ago |
| AudioShake Tasks API Hosted API for music, dialogue and effects separation, replacing AudioShake's legacy job API. | stem separation api | unknown | active | Checked 11d ago |
| Auphonic API API for automating Auphonic audio-processing workflows. | audio postproduction api | Austria | active | Checked 10d ago |
| Beat This! final0 Primary Beat This! checkpoint for beat and downbeat tracking without mandatory DBN postprocessing. | beat tracking model | unknown | active | Checked 11d ago |
| BlockMusic AI engine 블록뮤직 AI 엔진 Generative-music engine identified by Neutune as underlying its research and product development. | generative music engine | South Korea | active | Checked 10d ago |
| CLaMP (Sber music model) CLaMP Model described as converting a user text request into a representation consumed by SymFormer. | text music alignment model | Russia | active | Checked 10d ago |
| Cyanite Music Analysis API API for automated music tagging, emotional analysis, similarity search and catalog intelligence. | music analysis api | unknown | active | Checked 11d ago |
| Demucs v4 htdemucs Exact Hybrid Transformer Demucs v4 checkpoint separating 4 stems. | source separation model | unknown | active | Checked 11d ago |
| Demucs v4 htdemucs_6s Exact Hybrid Transformer Demucs v4 checkpoint separating 6 stems. | source separation model | unknown | active | Checked 11d ago |
| Demucs v4 htdemucs_ft Exact Hybrid Transformer Demucs v4 checkpoint separating 4 stems. | source separation model | unknown | active | Checked 11d ago |
| DiffRhythm Base 谛韵 Latent-diffusion full-song checkpoint targeting 1 minute 35 seconds output. | generative music model | unknown | active | Checked 11d ago |
| DiffRhythm Full 谛韵 Latent-diffusion full-song checkpoint targeting 4 minutes 45 seconds output. | generative music model | unknown | active | Checked 11d ago |
| DiffSinger Diffusion-based singing-voice synthesis model family with an official implementation and active OpenVPI maintenance. | singing voice synthesis model family | Mainland China, China | active | Checked 10d ago |
| EMSYNC Video-to-MIDI generation model aligning musical emotion and chord timing with video content and scene boundaries. | video conditioned symbolic music generation model | Porto, Portugal | active | Checked 10d ago |
| Eleven Music API Hosted REST API exposing Eleven Music v1 and v2 generation, planning, streaming and inpainting workflows. | hosted music generation api | unknown | active | Checked 11d ago |
| Eleven Music v1 First-generation Eleven Music model retained during a documented transition period. | generative music model | unknown | active | Checked 11d ago |
| Eleven Music v2 ElevenLabs' current text-to-music model with section planning, vocals, editing and inpainting. | generative music model | unknown | active | Checked 11d ago |
| FALL-E Gaudio Lab generative-audio engine for sound effects within Gaudio Studio. | generative sound effects | South Korea | active | Checked 10d ago |
| GSEP Gaudio Lab source-separation model for isolating vocals and instruments from mixed audio. | music source separation | South Korea | active | Checked 10d ago |
| GaMaDHaNi Hierarchical generative model of melodic vocal contours in Hindustani classical music. | music generation model | India | active | Checked 10d ago |
| Gemini Music Generation API Gemini API surface for Lyria 3 Clip, Lyria 3 Pro and Lyria RealTime music generation. | hosted music generation api | unknown | active | Checked 11d ago |
| Harmix API API for music catalog ingestion, metadata and AI-assisted music search workflows. | music catalog search api | Ukraine / global | active | Checked 10d ago |
| HeartCodec OSS 20260123 Released component of the HeartMuLa open music-foundation-model stack. | music audio codec | unknown | active | Checked 11d ago |
| HeartMuLa OSS 3B Released component of the HeartMuLa open music-foundation-model stack. | generative music model | unknown | active | Checked 11d ago |
| HeartTranscriptor OSS Released component of the HeartMuLa open music-foundation-model stack. | singing transcription model | unknown | active | Checked 11d ago |
| LAION Larger CLAP Music Contrastive audio-text model trained for music retrieval and classification. | audio text embedding model | unknown | active | Checked 11d ago |
| LAION Larger CLAP Music and Speech Contrastive audio-text model trained for music and speech retrieval and classification. | audio text embedding model | unknown | active | Checked 11d ago |
| Lyria 3 Google DeepMind music model for generating music and lyrics from prompts and other inputs. | generative music model | United Kingdom | active | Checked 11d ago |
| Lyria 3 Clip Preview Google's preview model for short music clips and loops. | generative music model | unknown | active | Checked 11d ago |
| Lyria 3 Pro Google music model for full songs and developer generation workflows. | music or audio model | unknown | active | Checked 11d ago |
| Lyria 3 Pro Preview Google's preview flagship for structured full-length songs. | generative music model | unknown | active | Checked 11d ago |
| Lyria RealTime Experimental Experimental Google model for continuously steerable instrumental-music streaming. | realtime generative music model | unknown | active | Checked 11d ago |
| MAGNeT Medium 10s Non-autoregressive MAGNeT checkpoint for 10 seconds of music. | generative music model | unknown | active | Checked 11d ago |
| MAGNeT Medium 30s Non-autoregressive MAGNeT checkpoint for 30 seconds of music. | generative music model | unknown | active | Checked 11d ago |
| MAGNeT Small 10s Non-autoregressive MAGNeT checkpoint for 10 seconds of music. | generative music model | unknown | active | Checked 11d ago |
| MAGNeT Small 30s Non-autoregressive MAGNeT checkpoint for 30 seconds of music. | generative music model | unknown | active | Checked 11d ago |
| MERT v1 330M Self-supervised acoustic-music representation checkpoint for downstream MIR tasks. | music embedding model | unknown | active | Checked 11d ago |
| MERT v1 95M Self-supervised acoustic-music representation checkpoint for downstream MIR tasks. | music embedding model | unknown | active | Checked 11d ago |
| Meta EnCodec 32 kHz (MusicGen checkpoint) Neural audio codec checkpoint trained for MusicGen tokenization. | music audio codec | unknown | active | Checked 11d ago |
| MiniMax Music 2.6 MiniMax model for original song, instrumental and cover generation exposed through product and API. | generative music model | Shanghai, Mainland China, China | active | Checked 10d ago |
| Moises-Light Resource-efficient band-split U-Net architecture for music source separation developed by Music AI researchers. | music source separation model | Brazil | active | Checked 10d ago |
| MuQ Large MSD Iter Released MuQ-family checkpoint for music representation and retrieval. | music embedding model | unknown | active | Checked 11d ago |
| MuQ-MuLan Large Released MuQ-family checkpoint for music representation and retrieval. | music embedding model | unknown | active | Checked 11d ago |
| Mubert AI Music API v3 REST API for text-to-music track generation, streaming, similarity and stem-level editing. | hosted music generation api | unknown | active | Checked 11d ago |
| Mureka O2 Current Mureka reasoning-oriented music-generation model released with V7.6. | music reasoning and generation model | Beijing, Mainland China, China | active | Checked 10d ago |
| Mureka V7.6 Current Mureka generative-music model released alongside O2. | generative music model | Beijing, Mainland China, China | active | Checked 10d ago |
| Music Flamingo 8B NVIDIA/UMD audio-language model for full-song understanding, reasoning and lyric transcription. | music understanding model | unknown | active | Checked 11d ago |
| MusicGen Large Exact MusicGen checkpoint for mono conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Medium Exact MusicGen checkpoint for mono conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Melody Exact MusicGen checkpoint for mono conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Melody Large Exact MusicGen checkpoint for mono conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Small Exact MusicGen checkpoint for mono conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Stereo Large Exact MusicGen checkpoint for stereo conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Stereo Medium Exact MusicGen checkpoint for stereo conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Stereo Melody Exact MusicGen checkpoint for stereo conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Stereo Melody Large Exact MusicGen checkpoint for stereo conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicGen Stereo Small Exact MusicGen checkpoint for stereo conditional music generation. | generative music model | unknown | active | Checked 11d ago |
| MusicLDM Text-to-music latent diffusion model with University of Surrey co-authorship. | text to music model | United Kingdom / United States research collaboration | active | Checked 10d ago |
| Mustango SUTD text-to-music model controllable through chords, beats, tempo and key instructions. | text to music model | Singapore | active | Checked 10d ago |
| N-Singer Non-autoregressive Korean singing-voice synthesis system described in the N-Singer paper. | korean singing voice synthesis | South Korea | active | Checked 10d ago |
| NVIDIA Fugatto 1 Research model for free-form audio synthesis and transformation from text and optional audio. | generative audio model | unknown | unknown | Checked 11d ago |
| Neural Capture Neural DSP machine-learning technology for learning the sound characteristics of amplifiers, cabinets and drive devices. | audio equipment modeling technology | Helsinki, Finland | active | Checked 10d ago |
| Notochord Low-latency probabilistic MIDI performance model designed for interactive improvisation and generation. | generative midi performance model | Reykjavík, Iceland | active | Checked 10d ago |
| Qinyue 琴乐音乐大模型 TME music model applied to AI song creation. | music generation model | Shenzhen, Mainland China, China | active | Checked 10d ago |
| Qinyun 琴韵音乐大模型 TME music model applied to AI singing and voice workflows. | singing voice model | Shenzhen, Mainland China, China | active | Checked 10d ago |
| RAVE Realtime Audio Variational autoEncoder for fast neural waveform synthesis and timbre transfer. | neural audio autoencoder | Paris, France | active | Checked 10d ago |
| RIMA AI 리마 AI Emotionwave engine for music analysis, generation, recommendation, and automated performance systems. | music analysis generation engine | South Korea | active | Checked 10d ago |
| Respeecher Space Real-time text-to-speech API operated by Respeecher. | realtime text to speech api | Ukraine / global | active | Checked 10d ago |
| Seed-Music ByteDance suite for controllable song generation, score-to-audio editing and zero-shot singing voice conversion. | controllable music generation system | Beijing, Mainland China, China | active | Checked 10d ago |
| SongGeneration 2 Large 4B LeVo 2 open checkpoint for multilingual song generation from text and audio prompts. | generative music model | unknown | active | Checked 11d ago |
| Spleeter 2stems Deezer pretrained TensorFlow separation model producing 2 stems. | source separation model | unknown | active | Checked 11d ago |
| Spleeter 4stems Deezer pretrained TensorFlow separation model producing 4 stems. | source separation model | unknown | active | Checked 11d ago |
| Spleeter 5stems Deezer pretrained TensorFlow separation model producing 5 stems. | source separation model | unknown | active | Checked 11d ago |
| Spotify Basic Pitch 0.4.0 Lightweight polyphonic audio-to-MIDI model with pitch-bend detection. | automatic music transcription model | unknown | active | Checked 11d ago |
| Stability AI Stable Audio 3 API Asynchronous hosted API for Stable Audio 3 text-to-audio, audio-to-audio and inpainting. | hosted audio generation api | unknown | active | Checked 11d ago |
| Stable Audio 3 Generative-audio model family with commercial and open-weight variants. | generative audio model | United Kingdom | active | Checked 11d ago |
| Stable Audio 3.0 Large Stable Audio 3 exact tier for low-latency high-volume generation, text-to-audio, audio-to-audio, inpainting. | generative audio model | unknown | active | Checked 11d ago |
| Stable Audio 3.0 Medium Stable Audio 3 exact tier for higher-musicality long-form generation. | generative audio model | unknown | active | Checked 11d ago |
| Stable Audio 3.0 Small Stable Audio 3 exact tier for music composition on-device. | generative audio model | unknown | active | Checked 11d ago |
| Stable Audio 3.0 Small SFX Stable Audio 3 exact tier for sound-effects generation on-device. | generative audio model | unknown | active | Checked 11d ago |
| Stable Audio Open 1.0 Open-weight latent-diffusion model for short text-to-audio generation. | generative audio model | unknown | active | Checked 11d ago |
| Suno v4.5-all Suno model exposed on the free tier at the snapshot date. | generative music model | unknown | active | Checked 11d ago |
| Suno v5.5 Suno's current personalized full-song generation model with custom-model and voice features. | generative music model | unknown | active | Checked 11d ago |
| SymFormer (Sber music model) SymFormer Automatic symbolic music generator producing notation-based compositions. | symbolic music generation model | Russia | active | Checked 10d ago |
| Tempolor 1.0 First documented Tempolor-series model for text-conditioned melody, rhythm and section structure generation. | controllable music generation model | Guangzhou, Mainland China, China | active | Checked 10d ago |
| Udio v1.5 Udio model version retained during the announced transition period. | music or audio model | unknown | active | Checked 11d ago |
| VOCALOID:AI AI synthesis engine used by VOCALOID6-compatible voicebanks. | singing voice synthesis engine | Japan | active | Checked 10d ago |
| Vobile AI Song Detector API OAuth-protected API classifying songs as AI-generated or human and optionally attributing a generator platform. | ai music detection api | unknown | active | Checked 11d ago |
| Voz Cantora Latam Experimental singing-voice model used in FUTURX research and a documented creative experiment. | experimental singing voice model | Latin America | active | Checked 10d ago |
| Whisper large-v3 Multilingual ASR and speech-to-English translation checkpoint. | speech recognition model | unknown | active | Checked 11d ago |
| Whisper large-v3-turbo Pruned four-decoder-layer ASR checkpoint optimized for speed. | speech recognition model | unknown | active | Checked 11d ago |
| YuE Stage 1 7B Anneal EN CoT 乐 Exact YuE checkpoint in the two-stage long-form song pipeline. | generative music model component | unknown | active | Checked 11d ago |
| YuE Stage 1 7B Anneal ZH CoT 乐 Exact YuE checkpoint in the two-stage long-form song pipeline. | generative music model component | unknown | active | Checked 11d ago |
| YuE Stage 2 1B General 乐 Exact YuE checkpoint in the two-stage long-form song pipeline. | generative music model component | unknown | active | Checked 11d ago |
| authio API REST API exposing authio AI-generated music detection. | ai music detection api | Saint-Denis, France | active | Checked 10d ago |
| maestro Beatoven.ai model for background music and sound-effect generation. | music or audio model | unknown | active | Checked 11d ago |
| music2latent Sony CSL Paris model for general-purpose latent audio compression. | latent audio compression model | Paris, France | active | Checked 10d ago |
| musicnn Pretrained musically motivated convolutional neural networks for music audio tagging and feature extraction. | music audio tagging model family | Barcelona, Catalonia, Spain | active | Checked 10d ago |
