MCP server for Spoken — fetch podcast transcripts as clean Markdown with real speaker names. Built for AI agents.
If you already use Spoken, the spoken mcp server is the piece that lets your assistant work with it directly. MCP server for Spoken — fetch podcast transcripts as clean Markdown with real speaker names. Built for AI agents.
Spoken is a transcript API that turns any published podcast into clean Markdown with real speaker names — not "Speaker 1." One API call returns named, timestamped text, ready for LLMs, RAG pipelines, summarizers, and search.
The toolset is worth reading before you wire it up, because it tells you what the integration is really for:
LLM — friendly overview: https://spoken.md/llms.txtsearch_podcasts — Find episodes by text or a pasted Spotify/YouTube URLlist_episodes — List a show's entire back-catalog from a podcast_idget_transcript — Fetch an episode's transcript as Markdown with real speaker namesget_balance — Check remaining creditsConfiguration is passed through the environment: SPOKEN_API_KEY. Treat anything key-shaped as a real credential — scope it to the minimum the server needs, and rotate it if it ever lands in a shared config.
Setup follows the usual MCP pattern — install or clone the server, register it in your client's configuration file, restart the client. The configuration blocks on this page cover the common clients.
Plenty of AI and media services servers cover similar ground. The differences that matter in practice are scope of access and how much setup stands between you and a working tool call. Spoken's toolset — LLM, search_podcasts, list_episodes and 2 more — is a fair guide to whether it matches your workflow. It is maintained by spokenmd; worth a glance at recent repository activity before you build anything load-bearing on it.
We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.
| Tool | What it does |
|---|---|
| LLM | friendly overview: https://spoken.md/llms.txt |
| search_podcasts | Find episodes by text or a pasted Spotify/YouTube URL |
| list_episodes | List a show's entire back-catalog from a podcast_id |
| get_transcript | Fetch an episode's transcript as Markdown with real speaker names |
| get_balance | Check remaining credits |
{
"mcpServers": {
"spoken": {
"command": "npx",
"args": ["-y", "spoken-mcp"],
"env": { "SPOKEN_API_KEY": "pt_your_key" }
}
}
}Configuration as documented by the project. Restart the client after saving.
| Variable | Description | Required |
|---|---|---|
| SPOKEN_API_KEY | Credential the server authenticates with. | Yes |
Build a programmable telecommunications stack for connecting telephony services with the Internet via a cloud-based utility.
Search built for AI, not humans — semantic web search that returns model-ready content, plus code context.
Answers, not links — delegate questions to Perplexity's search-grounded models and get cited responses back.
Give your assistant a voice — text-to-speech, voice cloning and audio tools from the ElevenLabs API.
Give your assistant a real code sandbox — isolated cloud VMs for actually running the code it writes.
The ML hub in your context window — search models, datasets, papers and run Spaces from the official server.