Fast and accurate web content extraction
Fast and accurate web content extraction. The trafilatura mcp server wraps that behind the Model Context Protocol, so an assistant can use it rather than through you.
const html = <html> <head><title>Example Article</title></head> <body> <nav>Home | About | Contact</nav> <article> <h1>Main Title</h1> This is the main content of the article. </article> <footer>Copyright 2024</footer> </body> </html>
The server ships on npm as trafilatura, so your MCP client can launch it on demand — there is no separate build step. Add the server block to your client's configuration, restart it, and the tools register themselves.
Among the knowledge and memory options, the useful question is rarely "what can it do" but "what does it cost you to run" — permissions, credentials, and how much of your context its toolset consumes. It is maintained by GitHub Actions; worth a glance at recent repository activity before you build anything load-bearing on it.
This entry was verified against Trafilatura's own documentation before publication; SyncDev keeps the directory reviewed rather than auto-generated.
{
"mcpServers": {
"trafilatura": {
"command": "npx",
"args": ["-y", "trafilatura"]
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
A knowledge graph your assistant keeps between sessions — entities, relations and observations that persist.
Kill hallucinated APIs — version-accurate, up-to-date library documentation injected straight into context.
Your workspace, on speaking terms with AI — search, read and write Notion pages and databases.
A structured scratchpad for hard problems — stepwise reasoning with revisions, branches and visible logic.
Symbol-level code navigation, refactoring and memory for coding agents — the IDE brain your assistant has been missing.
Chat with your second brain — search, read and write vault notes through the Local REST API.