Skip to content

phoenix

varize-phoenix-evals-v3.1.1 scope: arize-phoenix-evals Bugfix

This release fixes issues for SREs watching stability and regressions.

Published 13d Tracing
✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

agents ai-monitoring ai-observability aiengineering anthropic datasets
+10 more
evals langchain llamaindex llm-eval llm-evaluation llmops llms openai prompt-engineering smolagents

Summary

AI summary

Fixed evals bugs: correct per‑class weighted F‑score, count executor timeouts toward retry limit, and conditionally enable 0/1 label detection.

Changes in this release

Bugfix Medium

Counts AsyncExecutor timeouts against max_retries in evals.

Counts AsyncExecutor timeouts against max_retries in evals.

Source: llm_adapter@2026-07-14

Confidence: high

Bugfix Medium

Gates 0/1 positive_label auto-detection on default macro average in evals.

Gates 0/1 positive_label auto-detection on default macro average in evals.

Source: llm_adapter@2026-07-14

Confidence: high

Bugfix Medium

Computes macro/weighted F-score per class to match sklearn semantics in evals.

Computes macro/weighted F-score per class to match sklearn semantics in evals.

Source: llm_adapter@2026-07-14

Confidence: high

Full changelog

3.1.1 (2026-07-14)

Bug Fixes

  • evals: compute macro/weighted F-score per class to match sklearn semantics (#13740) (c0f6267)
  • evals: count AsyncExecutor timeouts against max_retries (#14361) (90eee08)
  • evals: gate 0/1 positive_label auto-detection on default macro average (#14012) (f96dbd9)

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track phoenix

Get notified when new releases ship.

Sign up free

About phoenix

AI Observability & Evaluation

All releases →

Related context

Earlier breaking changes

Beta — feedback welcome: [email protected]