Benchmark local LLM models — speed, quality & hardware fitness verdict from any MCP client
Connect Metrillm to Claude, Cursor or any other MCP client and it stops being a tab you switch to. Benchmark local LLM models — speed, quality & hardware fitness verdict from any MCP client. The metrillm mcp server is what makes that connection.
By default, production builds upload shared results to the official MetriLLM leaderboard (https://metrillm.dev).
Because this one is hosted, setup is mostly authentication — you point your client at the endpoint and approve access. Nothing runs on your machine, so there is no runtime to keep patched.
Configuration is passed through the environment: METRILLM_SUPABASE_URL, METRILLM_SUPABASE_ANON_KEY, METRILLM_PUBLIC_RESULT_BASE_URL, OLLAMA_HOST, LM_STUDIO_BASE_URL, LM_STUDIO_API_KEY. Treat anything key-shaped as a real credential — scope it to the minimum the server needs, and rotate it if it ever lands in a shared config.
Among the AI and media services options, the useful question is rarely "what can it do" but "what does it cost you to run" — permissions, credentials, and how much of your context its toolset consumes. It is maintained by MetriLLM; worth a glance at recent repository activity before you build anything load-bearing on it.
SyncDev reviews every entry in this directory against the project's own documentation before publishing, and revisits them as servers change.
{
"mcpServers": {
"metrillm": {
"command": "npx",
"args": ["-y", "metrillm"],
"env": {
"METRILLM_SUPABASE_URL": "your-value",
"METRILLM_SUPABASE_ANON_KEY": "your-value",
"METRILLM_PUBLIC_RESULT_BASE_URL": "your-value",
"OLLAMA_HOST": "your-value",
"LM_STUDIO_BASE_URL": "your-value",
"LM_STUDIO_API_KEY": "your-value"
}
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
| Variable | Description | Required |
|---|---|---|
| METRILLM_SUPABASE_URL | Endpoint or connection string the server talks to. | Yes |
| METRILLM_SUPABASE_ANON_KEY | Credential the server authenticates with. | Yes |
| METRILLM_PUBLIC_RESULT_BASE_URL | Endpoint or connection string the server talks to. | Yes |
| OLLAMA_HOST | Endpoint or connection string the server talks to. | Optional |
| LM_STUDIO_BASE_URL | Endpoint or connection string the server talks to. | Yes |
| LM_STUDIO_API_KEY | Credential the server authenticates with. | Yes |
Build a programmable telecommunications stack for connecting telephony services with the Internet via a cloud-based utility.
Search built for AI, not humans — semantic web search that returns model-ready content, plus code context.
Answers, not links — delegate questions to Perplexity's search-grounded models and get cited responses back.
Give your assistant a voice — text-to-speech, voice cloning and audio tools from the ElevenLabs API.
Give your assistant a real code sandbox — isolated cloud VMs for actually running the code it writes.
The ML hub in your context window — search models, datasets, papers and run Spaces from the official server.