Section 5 of 7 — 14 models
Audio, Voice & MusicAudio, Voice & Music
Song and instrumental generation, voice cloning and text-to-speech, and speech-to-text transcription.
- +Best-in-class vocals, capturing whispers, vibrato, and emotional nuance
- +Full song structure with proper verse, chorus, and bridge arrangement
- −Rap and spoken word still sound noticeably synthetic
- −No official public API yet; a partner API was only announced as exploratory in July 2026
- +Inpainting lets you regenerate one section without redoing the whole track
- +Stem separation and an official API for paid tiers
- −Public documentation of the v4 release is thin compared with Suno’s
- −Check licensing terms carefully before commercial use of generated music
- +Trained on licensed catalogs, giving strong legal safety for commercial use
- +Realistic voice cloning and text-to-speech with broad API access
- −Music composition quality trails Suno and Udio
- −Generation is slower and pricier than most competitors
- +Generates vocals with auto-written lyrics from text, image, or video prompts
- +High output quality for short-form music
- −Full-length songs need the Lyria 3 Pro tier; the base model makes short clips
- −Google has already moved on to Lyria 3.5 in its own music tools
- +Most affordable API-based music option available
- +Handles niche genre details well, with full commercial rights
- −Third-party API routes often cap clip length well below native limits
- −Much smaller brand recognition than Suno or Udio
- +Generates both music and sound effects, useful for production work
- +Audio inpainting for fine-tuning specific sections
- −No vocal generation
- −Commercial use is restricted by revenue thresholds
- +Robust, widely-used open speech-to-text across many languages and accents
- +Free and self-hostable, with a large surrounding tool ecosystem
- −No built-in speaker diarization out of the box
- −Architecture is aging relative to newer transcription models
- +Granular voice control, including emphasis, pitch, pacing, and pronunciation via IPA
- +Bundles voice cloning, dubbing, and translation alongside core text-to-speech
- −Voice library size and language coverage are less clearly documented than rivals
- −Free-tier limits and character caps aren't fully transparent
- +Large voice library with strong multilingual coverage
- +Good API access for developers building voice into products
- −Voice realism can vary noticeably across less common languages
- −Pricing tiers can get expensive at high usage volumes
- +Studio-quality, natural-sounding voices favored for corporate narration
- +Strong focus on brand-safe, licensed voice talent
- −Smaller voice selection than mass-market competitors
- −Positioned mainly at enterprise budgets, less accessible for casual users
- +Voice cloning built directly into a full audio and video editing workflow
- +Lets you edit spoken audio like text, moving words to reshape a recording
- −Requires recording your own voice samples to train a usable clone
- −Best value comes from Descript's editor, not the voice model alone
- +Focused on instrumental composition for film, games, and content creators
- +Lets users guide style and structure with more compositional control than most
- −No vocal generation
- −Less mainstream brand recognition than Suno or Udio
- +Extremely fast, one-click song creation aimed at total beginners
- +Built-in path to distribute finished tracks to streaming platforms
- −Creative control is shallow compared to Suno or Udio
- −Output quality trails the leading music generators
- +Deep integration with AWS, useful for developers already on that stack
- +Reliable, low-cost text-to-speech at scale for IVR and accessibility use
- −Voice realism trails newer generative voice platforms like ElevenLabs
- −Best suited to utilitarian use cases rather than expressive narration