Data Aggregator MCP Server

Find & fetch research datasets across Zenodo, DataCite, NCBI omics, and literature.

Remote serverstreamable-httpPython

What is the Data Aggregator MCP server?

Find & fetch research datasets across Zenodo, DataCite, NCBI omics, and literature. That is what the data aggregator mcp server brings to an AI assistant: the same capability, reachable through the Model Context Protocol rather than a separate app or dashboard.

The short version

search one query across 17 sourcesZenodo, DataCite (Dryad / Figshare / Dataverse / OSF / OpenNeuro / Mendeley), NCBI omics (GEO / SRA / BioProject), BioStudies (EBI, incl. ArrayExpress), literature (PubMed / OpenAIRE), HuggingFace datasets, DataONE (eco / environmental), OmicsDI (proteomics / metabolomics), DANDI (neurophysiology), CZ CELLxGENE (single-cell), OpenML (ML datasets), RCSB PDB (structures), UniProtKB (proteins), the GWAS Catalog, GBIF (biodiversity), data.gov (US federal open data), and NASA CMR (Earth science) — deduplicated

Getting it running

Because this one is hosted, setup is mostly authentication — you point your client at the endpoint and approve access. Nothing runs on your machine, so there is no runtime to keep patched.

The tools it exposes

The server publishes 1 tool. What each one is for:

  • Prompts — Three workflow prompts surface in clients (e.g. /mcp__data_aggregator__* in Claude Code):

What it needs from you

Configuration is passed through the environment: NCBI_API_KEY, DATA_GOV_API_KEY, DEMO_KEY, DATAVERSE_BASE_URL, EMBEDDING_API_KEY, LLM_API_KEY. Treat anything key-shaped as a real credential — scope it to the minimum the server needs, and rotate it if it ever lands in a shared config.

Things to watch

  • Your data travels to the provider's service, so the usual questions apply about what you send and what they retain.
  • Missing credentials fail quietly in some clients — if no tools show up, check the environment block first.
  • Keep per-call confirmation enabled while you learn its behaviour; it is the cheapest safeguard you have.

How it compares

This sits in the developer tooling group, where several servers overlap in what they claim to do but differ sharply once you actually set them up. Data Aggregator's toolset — Prompts — is a fair guide to whether it matches your workflow. It is maintained by musharna; worth a glance at recent repository activity before you build anything load-bearing on it.

We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.

Available tools

ToolWhat it does
PromptsThree workflow prompts surface in clients (e.g. /mcp__data_aggregator__* in Claude Code):

How to install the Data Aggregator MCP server

{
  "mcpServers": {
    "data-aggregator": {
      "command": "uvx",
      "args": ["data-aggregator-mcp"],
      "env": {
        "NCBI_API_KEY": "your-value",
        "DATA_GOV_API_KEY": "your-value",
        "DEMO_KEY": "your-value",
        "DATAVERSE_BASE_URL": "your-value",
        "EMBEDDING_API_KEY": "your-value",
        "LLM_API_KEY": "your-value"
      }
    }
  }
}

Add to claude_desktop_config.json, then restart Claude Desktop.

Configuration

VariableDescriptionRequired
NCBI_API_KEYCredential the server authenticates with.Yes
DATA_GOV_API_KEYCredential the server authenticates with.Yes
DEMO_KEYCredential the server authenticates with.Yes
DATAVERSE_BASE_URLEndpoint or connection string the server talks to.Yes
EMBEDDING_API_KEYCredential the server authenticates with.Yes
LLM_API_KEYCredential the server authenticates with.Yes

Example prompts to try

  • Use Data Aggregator to Prompts.

Frequently asked questions

It connects Data Aggregator to MCP-compatible AI assistants such as Claude and Cursor, exposing 1 tool (Prompts) that the assistant can call on your behalf. Instead of copying data back and forth by hand, the assistant works with Data Aggregator directly.