This release adds 1 notable feature for engineering teams evaluating rollout.
✓ No known CVEs patched in this version
Topics
+5 more
Summary
AI summaryAdded per‑evaluation spend budget for LLM‑as‑judge evaluators.
Changes in this release
| Type | Severity | Summary | CVE |
|---|---|---|---|
| Feature | Medium |
Add per-evaluation spend budget for LLM-as-judge evaluators. Add per-evaluation spend budget for LLM-as-judge evaluators. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Feature | Low |
Enable cost-api auth in dev-runner platform mode. Enable cost-api auth in dev-runner platform mode. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Dependency | Low |
Sync provider model definitions across backend and frontend. Sync provider model definitions across backend and frontend. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Dependency | Low |
Automatically update OpenAPI spec and Fern code. Automatically update OpenAPI spec and Fern code. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Performance | Medium |
Optimize full-scan audit Slice-1 remediations (R4, R6, R7, R8). Optimize full-scan audit Slice-1 remediations (R4, R6, R7, R8). Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Performance | Medium |
Add project/trace-id scope dataset-item and optimize enrichment scans. Add project/trace-id scope dataset-item and optimize enrichment scans. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Bugfix | Medium |
Fix cheap existence probe for Logs empty state. Fix cheap existence probe for Logs empty state. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Bugfix | Low |
Retry model-picker click to fix Playground flake. Retry model-picker click to fix Playground flake. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Bugfix | Low |
Add copy for permission_denied run failure diagnostics. Add copy for permission_denied run failure diagnostics. Source: llm_adapter@2026-07-15 Confidence: high |
— |
| Bugfix | Low |
Update model prices file. Update model prices file. Source: llm_adapter@2026-07-15 Confidence: high |
— |
Full changelog
What's Changed
- [OPIK-6884] [BE] perf: full-scan audit Slice-1 remediations (R4, R6, R7, R8) by @thiagohora in https://github.com/comet-ml/opik/pull/7404
- [NA] [QA] test(e2e): retry model-picker click to fix Playground flake by @AndreiCautisanu in https://github.com/comet-ml/opik/pull/7401
- [OPIK-6993] [BE] Capture Claude Agent SDK / Claude Code content in Opik OTEL traces by @jverre in https://github.com/comet-ml/opik/pull/7319
- [NA] [FE] Diagnostics: copy for permission_denied run failure by @aadereiko in https://github.com/comet-ml/opik/pull/7407
- [OPIK-7249] [BE][FE] fix: cheap existence probe for Logs empty state by @YarivHashaiComet in https://github.com/comet-ml/opik/pull/7378
- [NA] [DOCS] Add Diagnostics documentation page by @jverre in https://github.com/comet-ml/opik/pull/7417
- [OPIK-6884] [BE] perf: project/trace-id scope dataset-item + optimization enrichment scans by @thiagohora in https://github.com/comet-ml/opik/pull/7412
- [OPIK-7245] dev-runner: enable cost-api auth in platform mode by @LifeXplorer in https://github.com/comet-ml/opik/pull/7415
- [NA] [BE] Update model prices file by @CometActions in https://github.com/comet-ml/opik/pull/7425
- [NA] [BE][FE] chore: sync provider model definitions by @CometActions in https://github.com/comet-ml/opik/pull/7426
- [OPIK-6995] [FS] feat: add per-evaluation spend budget for LLM-as-judge evaluators by @alexkuzmik in https://github.com/comet-ml/opik/pull/7312
- [NA] [SDK] [DOCS] Update automatically OpenAPI spec and Fern code by @CometActions in https://github.com/comet-ml/opik/pull/7424
Full Changelog: https://github.com/comet-ml/opik/compare/2.1.21...2.1.22
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Related context
Related tools
Beta — feedback welcome: [email protected]