Find & fetch research datasets across Zenodo, DataCite, NCBI omics, and literature.
Find & fetch research datasets across Zenodo, DataCite, NCBI omics, and literature. That is what the data aggregator mcp server brings to an AI assistant: the same capability, reachable through the Model Context Protocol rather than a separate app or dashboard.
search one query across 17 sources — Zenodo, DataCite (Dryad / Figshare / Dataverse / OSF / OpenNeuro / Mendeley), NCBI omics (GEO / SRA / BioProject), BioStudies (EBI, incl. ArrayExpress), literature (PubMed / OpenAIRE), HuggingFace datasets, DataONE (eco / environmental), OmicsDI (proteomics / metabolomics), DANDI (neurophysiology), CZ CELLxGENE (single-cell), OpenML (ML datasets), RCSB PDB (structures), UniProtKB (proteins), the GWAS Catalog, GBIF (biodiversity), data.gov (US federal open data), and NASA CMR (Earth science) — deduplicated
Because this one is hosted, setup is mostly authentication — you point your client at the endpoint and approve access. Nothing runs on your machine, so there is no runtime to keep patched.
The server publishes 1 tool. What each one is for:
Prompts — Three workflow prompts surface in clients (e.g. /mcp__data_aggregator__* in Claude Code):Configuration is passed through the environment: NCBI_API_KEY, DATA_GOV_API_KEY, DEMO_KEY, DATAVERSE_BASE_URL, EMBEDDING_API_KEY, LLM_API_KEY. Treat anything key-shaped as a real credential — scope it to the minimum the server needs, and rotate it if it ever lands in a shared config.
This sits in the developer tooling group, where several servers overlap in what they claim to do but differ sharply once you actually set them up. Data Aggregator's toolset — Prompts — is a fair guide to whether it matches your workflow. It is maintained by musharna; worth a glance at recent repository activity before you build anything load-bearing on it.
We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.
| Tool | What it does |
|---|---|
| Prompts | Three workflow prompts surface in clients (e.g. /mcp__data_aggregator__* in Claude Code): |
{
"mcpServers": {
"data-aggregator": {
"command": "uvx",
"args": ["data-aggregator-mcp"],
"env": {
"NCBI_API_KEY": "your-value",
"DATA_GOV_API_KEY": "your-value",
"DEMO_KEY": "your-value",
"DATAVERSE_BASE_URL": "your-value",
"EMBEDDING_API_KEY": "your-value",
"LLM_API_KEY": "your-value"
}
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
| Variable | Description | Required |
|---|---|---|
| NCBI_API_KEY | Credential the server authenticates with. | Yes |
| DATA_GOV_API_KEY | Credential the server authenticates with. | Yes |
| DEMO_KEY | Credential the server authenticates with. | Yes |
| DATAVERSE_BASE_URL | Endpoint or connection string the server talks to. | Yes |
| EMBEDDING_API_KEY | Credential the server authenticates with. | Yes |
| LLM_API_KEY | Credential the server authenticates with. | Yes |
Kill hallucinated APIs — version-accurate, up-to-date library documentation injected straight into context.
Microsoft's official browser automation server — drive a real browser through the accessibility tree, no screenshots needed.
GitHub's official server — repos, issues, pull requests, Actions and code security, straight from your assistant.
Issue tracking at the speed of conversation — Linear's official hosted server with OAuth and zero install.
Local repository surgery — status, diffs, commits, branches and history for any repo on disk.
Timezone sanity for AI — current time anywhere and correct conversions, without the model doing date math.