Skip to content

Rubric

LLM Frameworks

Agent behavior testing for LLM apps – verify which tools are called, their arguments, trace quality, latency and more to catch regressions in CI.

Python Latest v0.2.0 · 1mo ago Security brief →

Features

  • Validate tool calls (expected, forbidden, order)
  • Assess trace quality (no loops, step budget adherence)
  • Measure latency and cost budgets
  • Run LLM‑based or traditional string metrics on outputs

Recent releases

View all 2 releases →
No immediate action
v0.2.0 Breaking risk

LangGraph capture + regression diffing

No immediate action
v0.1.1 Breaking risk

overall_score rename + CLI fixes

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
13
Forks
1
Languages
Python Makefile

Install & Platforms

Install via
pip

Similar tools

Beta — feedback welcome: [email protected]