Query Spark SQL clusters via Thrift/HiveServer2. Works with Spark, EMR, Hive, Impala.
Connect Spark SQL to Claude, Cursor or any other MCP client and it stops being a tab you switch to. Query Spark SQL clusters via Thrift/HiveServer2. Works with Spark, EMR, Hive, Impala. The spark sql mcp server is what makes that connection.
An MCP server that enables AI assistants to query Spark SQL clusters via the Thrift/HiveServer2 protocol.
The toolset is worth reading before you wire it up, because it tells you what the integration is really for:
list_databases — List all available databaseslist_tables — List tables in a databasedescribe_table — Get table schema (columns, types)execute_query — Run read-only SQL queries with formatted resultsLDAP — The LDAP tool exposed by this serverKerberos — The Kerberos tool exposed by this serverConfiguration is passed through the environment: SPARK_HOST, SPARK_PORT, SPARK_AUTH, SPARK_DATABASE, SPARK_USERNAME, SPARK_PASSWORD, SPARK_KERBEROS_SERVICE_NAME. Treat anything key-shaped as a real credential — scope it to the minimum the server needs, and rotate it if it ever lands in a shared config.
The server ships on PyPI as spark-sql-mcp-server, so your MCP client can launch it on demand — there is no separate build step. Add the server block to your client's configuration, restart it, and the tools register themselves.
Plenty of database access servers cover similar ground. The differences that matter in practice are scope of access and how much setup stands between you and a working tool call. Spark SQL's toolset — list_databases, list_tables, describe_table and 3 more — is a fair guide to whether it matches your workflow. It is maintained by aidancorrell; worth a glance at recent repository activity before you build anything load-bearing on it.
We check each listing at SyncDev against the project's documentation before it goes live — if something here drifts out of date, it is a bug worth reporting.
| Tool | What it does |
|---|---|
| list_databases | List all available databases |
| list_tables | List tables in a database |
| describe_table | Get table schema (columns, types) |
| execute_query | Run read-only SQL queries with formatted results |
| LDAP | The LDAP tool exposed by this server. |
| Kerberos | The Kerberos tool exposed by this server. |
{
"mcpServers": {
"spark-sql": {
"command": "uvx",
"args": ["spark-sql-mcp-server"],
"env": {
"SPARK_HOST": "your-value",
"SPARK_PORT": "your-value",
"SPARK_AUTH": "your-value",
"SPARK_DATABASE": "your-value",
"SPARK_USERNAME": "your-value",
"SPARK_PASSWORD": "your-value",
"SPARK_KERBEROS_SERVICE_NAME": "your-value"
}
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
| Variable | Description | Required |
|---|---|---|
| SPARK_HOST | Endpoint or connection string the server talks to. | Optional |
| SPARK_PORT | Configuration value read at startup. | Optional |
| SPARK_AUTH | Configuration value read at startup. | Optional |
| SPARK_DATABASE | Configuration value read at startup. | Optional |
| SPARK_USERNAME | Configuration value read at startup. | Optional |
| SPARK_PASSWORD | Configuration value read at startup. | Optional |
| SPARK_KERBEROS_SERVICE_NAME | Configuration value read at startup. | Optional |
Read-only SQL access to Postgres — let your assistant inspect schemas and answer questions from real data.
Manage your whole Supabase project in conversation — database, auth, storage, Edge Functions and branches.
Query, modify and analyse local SQLite databases in conversation — the fastest way to chat with a data file.
Metabase ships its own MCP endpoint — search your BI content, build and run queries, and save questions and dashboards without leaving the chat.
Official MongoDB server covering data, schemas and Atlas management — from find queries to spinning up clusters.
Serverless Postgres with database branching — point your assistant at Neon and let it work on disposable copies.