This release adds 2 notable features for engineering teams evaluating rollout.
✓ No known CVEs patched in this version
Topics
+12 more
Summary
AI summaryAdded per-arm worker model support and archetype_facts env toggle in the effectiveness runner.
Full changelog
Model-tier arms in the effectiveness eval (roadmap #5). All existing lift /
no-lift evidence is sonnet-worker evidence; this makes the instrument that every
model-era decision depends on able to measure a stronger worker. Local-only eval
harness — no production hook, no user surface.
Added
- Per-arm worker model in the effectiveness runner:
--arm-model shadow=opus,enforce=fablespawns each arm's sessions on its own model (arms
not named fall back to--model). A paired toggle arm inherits its base arm's
model so the A/B isolates the feature, not the model. The effective model is
recorded per cell and inrun.json's newarm_modelsmap, and
compare_to_baselineis now model-aware: a legacy flat baseline answers only
the sonnet arm, a model-keyed baseline is matched per arm's model, so a
stronger-model arm never regresses against a sonnet baseline. archetype_facts(CHAMELEON_ARCHETYPE_FACTS) added to the eval's env-toggle
set, so the per-edit archetype-facts directive can be A/B'd like the other
default-on injection features (--toggle archetype_facts).
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About Chameleon
All releases →Beta — feedback welcome: [email protected]