Give your assistant a voice — text-to-speech, voice cloning and audio tools from the ElevenLabs API.
The ElevenLabs MCP server turns text conversations into audio production sessions. Backed by the same API that powers most of the convincing AI voices you hear online, it lets your assistant generate speech in hundreds of voices and dozens of languages, design entirely new voices from text descriptions, clone voices from samples, transcribe audio, and even isolate speech from noisy recordings.
The conversational framing genuinely fits this domain. Voiceover work is iterative — "warmer, slower, more conviction on the second line" — and iterating through an assistant that regenerates the clip in seconds beats round-tripping through a studio interface. Scripts written and voiced in the same conversation collapse a whole workflow: draft the ad copy, hear it, revise both words and delivery, export.
Practical touches worth knowing: output files land in your Desktop/ElevenLabs folder by default, voice-design tools let you conjure "a gravelly narrator with a slight Irish accent" without auditioning a library, and transcription closes the loop for podcast and interview workflows. Costs follow your ElevenLabs plan's character quota — generation is metered, so bulk audiobook ambitions need a paid tier, while the free tier covers experimentation.
Setup is one API key via uvx. If your work touches spoken audio at all — product demos, localised voiceovers, podcast tooling, accessibility — this is the most capable audio server in the MCP ecosystem, from the company that set the current bar for synthetic speech.
voice_id through the rest of a script keeps the character consistent instead of drifting line to line.Write, voice, critique and re-voice marketing or video copy iteratively.
The same script rendered across languages and voices for global content.
Transcribe episodes, clean noisy recordings, generate intros.
| Tool | What it does |
|---|---|
| text_to_speech | Generate speech from text in a chosen voice |
| search_voices / get_voice | Browse and inspect the voice library |
| voice_clone | Clone a voice from audio samples |
| text_to_voice | Design a new voice from a text description |
| speech_to_text | Transcribe audio files |
| isolate_audio | Separate speech from background noise |
| play_audio | Play generated audio locally |
{
"mcpServers": {
"elevenlabs": {
"command": "uvx",
"args": ["elevenlabs-mcp"],
"env": { "ELEVENLABS_API_KEY": "your-api-key" }
}
}
}An ElevenLabs account and API key; character quotas follow your plan. Python with uv.
| Variable | Description | Required |
|---|---|---|
| ELEVENLABS_API_KEY | API key from the ElevenLabs dashboard | Yes |
Search built for AI, not humans — semantic web search that returns model-ready content, plus code context.
Answers, not links — delegate questions to Perplexity's search-grounded models and get cited responses back.
Give your assistant a real code sandbox — isolated cloud VMs for actually running the code it writes.
The ML hub in your context window — search models, datasets, papers and run Spaces from the official server.
Thousands of open models on tap — run image, video and audio generation through Replicate's hosted API.
Pull live Ahrefs backlink, keyword and site-audit data straight into your AI assistant.