npcpy is the premier open-source library for integrating LLMs and Agents into python systems.
If you already use Npcpy, the npcpy mcp server is the piece that lets your assistant work with it directly. npcpy is the premier open-source library for integrating LLMs and Agents into python systems.
npcpy is a library that provides key primitives for research and development with multimodal language models, agentic AI, and knowledge graphs. Its flexible framework makes it easy to engineer powerful AI applications with support for local (ollama, llama.cpp, omlx, LM Studio) and cloud providers. Build multi-agent teams and simplify context engineering through the NPC Context-Agent-Tool data layer which ensures compliance through software rather than prompts. bash pip install npcpy
The toolset is worth reading before you wire it up, because it tells you what the integration is really for:
ToolAgent — Attach custom tools to a ToolAgent. Here is an example which lets an agent generate images, fine-tune diffusion models, and then use the fine-tunedStreaming — response = get_llm_response("Explain quantum entanglement.", model='qwen3.5:2b', provider='ollama', stream=True) for chunk in response['response']npcpy on PyPI is all you need. Most clients run it directly, so configuration is a few lines and a restart.
Plenty of developer tooling servers cover similar ground. The differences that matter in practice are scope of access and how much setup stands between you and a working tool call. Npcpy's toolset — ToolAgent, Streaming — is a fair guide to whether it matches your workflow. It is maintained by Christopher Agostino; worth a glance at recent repository activity before you build anything load-bearing on it.
We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.
| Tool | What it does |
|---|---|
| ToolAgent | Attach custom tools to a ToolAgent. Here is an example which lets an agent generate images, fine-tune diffusion models, and then use the fine-tuned models for generation. |
| Streaming | response = get_llm_response("Explain quantum entanglement.", model='qwen3.5:2b', provider='ollama', stream=True) for chunk in response['response']: content, _, _ = parse_stream_chunk(chunk, provider='ollama') if content: |
{
"mcpServers": {
"npcpy": {
"command": "uvx",
"args": ["npcpy"]
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
Kill hallucinated APIs — version-accurate, up-to-date library documentation injected straight into context.
Microsoft's official browser automation server — drive a real browser through the accessibility tree, no screenshots needed.
GitHub's official server — repos, issues, pull requests, Actions and code security, straight from your assistant.
Issue tracking at the speed of conversation — Linear's official hosted server with OAuth and zero install.
Local repository surgery — status, diffs, commits, branches and history for any repo on disk.
Timezone sanity for AI — current time anywhere and correct conversions, without the model doing date math.