Provides text-to-speech and speech-to-text capabilities using Windows' built-in speech services.
If you already use MS, the ms mcp server is the piece that lets your assistant work with it directly. Provides text-to-speech and speech-to-text capabilities using Windows' built-in speech services.
A Model Context Protocol (MCP) server that provides text-to-speech and speech-to-text capabilities using Windows' built-in speech services. This server leverages the native Windows Speech API (SAPI) through PowerShell commands, eliminating the need for external APIs or services.
The toolset is worth reading before you wire it up, because it tells you what the integration is really for:
text_to_speech — Parameters: - text (required): The text to convert to speech - voice (optional): The voice to use (e.g., "Microsoft David Desktop") - speedspeech_to_text — Records audio and converts it to text using Windows Speech RecognitionSetup follows the usual MCP pattern — install or clone the server, register it in your client's configuration file, restart the client.
Plenty of AI and media services servers cover similar ground. The differences that matter in practice are scope of access and how much setup stands between you and a working tool call. MS's toolset — text_to_speech, speech_to_text — is a fair guide to whether it matches your workflow. It is maintained by ExpressionsBot; worth a glance at recent repository activity before you build anything load-bearing on it.
We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.
| Tool | What it does |
|---|---|
| text_to_speech | Parameters: - text (required): The text to convert to speech - voice (optional): The voice to use (e.g., "Microsoft David Desktop") - speed (optional): Speech rate from 0.5 to 2.0 (default: 1.0) |
| speech_to_text | Records audio and converts it to text using Windows Speech Recognition. |
Build a programmable telecommunications stack for connecting telephony services with the Internet via a cloud-based utility.
Search built for AI, not humans — semantic web search that returns model-ready content, plus code context.
Answers, not links — delegate questions to Perplexity's search-grounded models and get cited responses back.
Give your assistant a voice — text-to-speech, voice cloning and audio tools from the ElevenLabs API.
Give your assistant a real code sandbox — isolated cloud VMs for actually running the code it writes.
The ML hub in your context window — search models, datasets, papers and run Spaces from the official server.