Skip to content

hidai25/eval-view

Developer Productivity

Regression testing framework for AI agents. Save golden baselines, detect behavioral drift, and block regressions in CI. Works with LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.

Python Latest v0.8.1 · 10h ago Security brief →

Features

  • Snapshot testing for AI agents – records tool calls, order, and output
  • Detects regressions or unexpected changes in agent behavior without manual assertions
  • Works offline; optional LLM judge for output‑quality scoring
  • CI integration to block regression merges
  • Supports multiple frameworks (LangGraph, CrewAI, OpenAI, Claude, Mistral, Ollama, MCP) and any HTTP API

Recent releases

View all 33 releases →
No immediate action
v0.8.1 Mixed

Watch fix + PEP561 + check --watch

No immediate action
v0.8.0 New feature

Cassettes + schedule cron

No immediate action
v0.7.1 Breaking risk

TOML test cases + CSV log import

No immediate action
v0.7.0 New feature

Aider CLI adapter

No immediate action
v0.6.2 New feature

Closed-model drift detection

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
123
Forks
21
Languages
Python Makefile TypeScript
Downloads/week
17 ↑5%
NPM Maintainers
1
Contributors
13
TypeScript
Types included ✓

Install & Platforms

Install via
pip

Alternative to

Langfuse LangSmith promptfoo DeepEval Braintrust

Beta — feedback welcome: [email protected]