Evalview MCP Server

Regression testing for AI agents. Golden baselines, CI/CD, LangGraph, CrewAI, OpenAI, Claude.

Local serverstdioPython

What is the Evalview MCP server?

Regression testing for AI agents. Golden baselines, CI/CD, LangGraph, CrewAI, OpenAI, Claude. The evalview mcp server wraps that behind the Model Context Protocol, so an assistant can use it through 4 defined tools rather than through you.

What it actually does

Record what your agent does today. Get told when it silently changes.

Its toolset

Everything the assistant can do here goes through one of these:

  • uses — actions/checkout@v4
  • Setup — Record current behavior
  • Catches — Any drift from baseline
  • Non-determinism — Multi-variant baselines (up to 5 valid paths)

Configuration

You will need one environment variable: OPENAI_API_KEY. The server will not start without them, which is usually why the tools fail to appear on a first run. Keep credentials in your client's env block or a secrets manager rather than in a file you might commit.

Adding it to your client

Installation goes through your MCP client rather than a global install: point it at evalview on PyPI and it is fetched when the client starts. The copy-paste blocks for Claude Desktop, Claude Code and Cursor are further down this page.

When to reach for it

Plenty of payments and commerce servers cover similar ground. The differences that matter in practice are scope of access and how much setup stands between you and a working tool call. Evalview's toolset — uses, Setup, Catches and 1 more — is a fair guide to whether it matches your workflow. It is maintained by hidai25; worth a glance at recent repository activity before you build anything load-bearing on it.

SyncDev reviews every entry in this directory against the project's own documentation before publishing, and revisits them as servers change.

Caveats

  • It runs with your machine's permissions. That is convenient and also the reason to think about what you point it at before you approve a tool call.
  • Missing credentials fail quietly in some clients — if no tools show up, check the environment block first.
  • MCP clients confirm each tool call by default. Leave that on until you have watched what the evalview mcp server does with a few real requests.

Available tools

ToolWhat it does
usesactions/checkout@v4
SetupRecord current behavior
CatchesAny drift from baseline
Non-determinismMulti-variant baselines (up to 5 valid paths)

How to install the Evalview MCP server

{
  "mcpServers": {
    "evalview": {
      "command": "uvx",
      "args": ["evalview"],
      "env": {
        "OPENAI_API_KEY": "your-value"
      }
    }
  }
}

Add to claude_desktop_config.json, then restart Claude Desktop.

Configuration

VariableDescriptionRequired
OPENAI_API_KEYCredential the server authenticates with.Yes

Example prompts to try

  • Use Evalview to uses.
  • Use Evalview to Setup.
  • Use Evalview to Catches.

Frequently asked questions

It connects Evalview to MCP-compatible AI assistants such as Claude and Cursor, exposing 4 tools (uses, Setup, Catches, and more) that the assistant can call on your behalf. Instead of copying data back and forth by hand, the assistant works with Evalview directly.