Running local LLMs means choosing between a strong reasoning model and a fast coding model. You can't load both on a single machine. Manually
Running local LLMs means choosing between a strong reasoning model and a fast coding model. You can't load both on a single machine. Manually swapping models kills your conversation context and flow. The mcp llama swap mcp server wraps that behind the Model Context Protocol, so an assistant can use it through 5 defined tools rather than through you.
mcp-llama-swap on PyPI is all you need. Most clients run it directly, so configuration is a few lines and a restart.
Everything the assistant can do here goes through one of these:
list_models — Lists all configured models with load status and current modeget_current_model — Returns the alias of the currently loaded modelswap_model — Unloads current model, loads the specified one, waits for health checkcreate_model_config — Generates a new launchd plist (macOS) or systemd unit (Linux) for a modelPrerequisites — The Prerequisites tool exposed by this serverYou will need 4 environment variables: LLAMA_SWAP_CONFIG, ANTHROPIC_BASE_URL, ANTHROPIC_API_KEY, ANTHROPIC_MODEL. The server will not start without them, which is usually why the tools fail to appear on a first run. Keep credentials in your client's env block or a secrets manager rather than in a file you might commit.
This sits in the developer tooling group, where several servers overlap in what they claim to do but differ sharply once you actually set them up. MCP Llama Swap's toolset — list_models, get_current_model, swap_model and 2 more — is a fair guide to whether it matches your workflow. It is maintained by oussama-kh; worth a glance at recent repository activity before you build anything load-bearing on it.
We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.
| Tool | What it does |
|---|---|
| list_models | Lists all configured models with load status and current mode |
| get_current_model | Returns the alias of the currently loaded model |
| swap_model | Unloads current model, loads the specified one, waits for health check |
| create_model_config | Generates a new launchd plist (macOS) or systemd unit (Linux) for a model |
| Prerequisites | The Prerequisites tool exposed by this server. |
{
"mcpServers": {
"llama-swap": {
"command": "uvx",
"args": ["mcp-llama-swap"],
"env": {
"LLAMA_SWAP_CONFIG": "your-value",
"ANTHROPIC_BASE_URL": "your-value",
"ANTHROPIC_API_KEY": "your-value",
"ANTHROPIC_MODEL": "your-value"
}
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
| Variable | Description | Required |
|---|---|---|
| LLAMA_SWAP_CONFIG | Configuration value read at startup. | Optional |
| ANTHROPIC_BASE_URL | Endpoint or connection string the server talks to. | Yes |
| ANTHROPIC_API_KEY | Credential the server authenticates with. | Yes |
| ANTHROPIC_MODEL | Configuration value read at startup. | Optional |
Kill hallucinated APIs — version-accurate, up-to-date library documentation injected straight into context.
Microsoft's official browser automation server — drive a real browser through the accessibility tree, no screenshots needed.
GitHub's official server — repos, issues, pull requests, Actions and code security, straight from your assistant.
Issue tracking at the speed of conversation — Linear's official hosted server with OAuth and zero install.
Local repository surgery — status, diffs, commits, branches and history for any repo on disk.
Timezone sanity for AI — current time anywhere and correct conversions, without the model doing date math.