Skip to content

Cortex

v1.6.0 Breaking

This release includes breaking changes for platform teams planning a safe upgrade.

Published 20h MCP Developer Tools
✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai-agents docker fastapi karpathy-inspired karpathy-llm-wiki knowledge-base
+8 more
knowledge-graph knowledge-management llm llm-wiki markdown mcp python qmd

Summary

AI summary

Added a reproducible retrieval benchmark harness, published LongMemEval‑S results at 45.0% accuracy, and native Metal‑GPU QMD support on macOS.

Full changelog

v1.6.0

Released 2026-07-27.

Benchmarks release: a reproducible retrieval-evaluation harness, the first published LongMemEval-S result, and native (Metal GPU) QMD on macOS.

Added

  • Retrieval benchmark harness (evals/): reproducible, auditable evaluation of Cortex retrieval end-to-end through the public REST API — ingest → compile → search → answer → judge, exactly as a user deployment runs it. Design principles:

    • Pinned configs: each run is a YAML file pinning dataset URL + SHA-256, search mode, answer model, judge model, and seed; a hash mismatch aborts the run.
    • Per-question vault isolation: the vault is wiped and QMD synchronously reindexed between questions, so retrieval only ever sees that question's haystack.
    • Honest judging: LLM-as-judge with strict yes/no verdict parsing (unparseable replies count as errors, never guesses), judge model enforced ≠ answer model, and the official LongMemEval per-type grading prompts verbatim (temporal off-by-one tolerance, preference rubric, knowledge-update semantics, abstention handling).
    • Auditable artifacts: per-question JSONL traces (question, retrieved context, answer, verdict, latency, recall@k from dataset evidence labels, per-phase token usage) streamed crash-safe during the run.
    • Blind validation: identifier-stripped sample packets for human review plus judge–human agreement (percent and Cohen's kappa).
  • First published benchmark result — LongMemEval-S, hybrid search: 45.0% overall accuracy, 65.6% evidence recall@8 over all 500 questions, zero judge errors (answer model anthropic/claude-sonnet-4, judge openai/gpt-4o). Per-category table, methodology, and caveats: docs/benchmarks/README.md. Reproduce with:

    uv run python -m evals.run_longmemeval --config evals/configs/longmemeval-s-hybrid.yaml
    
  • Native QMD mode (./scripts/start.sh --native-qmd / ./scripts/stop.sh): runs QMD directly on the host instead of in Docker. On Apple-silicon Macs this gives QMD Metal-GPU embedding — Docker Desktop cannot access the GPU, which made embedding-heavy workloads time out under CPU emulation. stop.sh detects and stops a native QMD automatically via its PID file.

Fixed

  • Trace files with Unicode line separators: read_traces splits records on newline only. Model answers containing U+2028/U+2029 (written unescaped with ensure_ascii=False) previously corrupted JSONL parsing via splitlines().

Upgrade notes

  • No breaking changes. The evals/ package is dev-only tooling and is not part of the published PyPI package's runtime surface.
  • Benchmark runs require LLM_API_KEY (and optionally LLM_BASE_URL) for the answer and judge models named in the config; Cortex itself still runs fine with LLM_BACKEND=none.

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track Cortex

Get notified when new releases ship.

Sign up free

About Cortex

All releases →

Related context

Earlier breaking changes

  • v1.2.0 Changes default search mode to hybrid from keyword.

Beta — feedback welcome: [email protected]