Nightly eval REGRESSION: benchmark 'contract_roman_numeral' passed ALL trials in the previous run but failed BOTH trials on 2026-07-19.
Error category: [compile_error]
Model: opencode-qwen3-5-35b-a3b-mxfp8 (local, tiers:smoke,core)
Results: /tmp/nightly_eval_20260719 (prev: /tmp/nightly_eval_20260718_rag_on/agent)
It was solid last night, so this is a fresh solid->broken regression — investigate.
Binary info (auto-attached):
ailang version: v0.29.2-421-g81a45f2d8
git commit: 81a45f2
Reported by: nightly-eval via ailang messages
Nightly eval REGRESSION: benchmark 'contract_roman_numeral' passed ALL trials in the previous run but failed BOTH trials on 2026-07-19.
Error category: [compile_error]
Model: opencode-qwen3-5-35b-a3b-mxfp8 (local, tiers:smoke,core)
Results: /tmp/nightly_eval_20260719 (prev: /tmp/nightly_eval_20260718_rag_on/agent)
It was solid last night, so this is a fresh solid->broken regression — investigate.
Binary info (auto-attached):
ailang version: v0.29.2-421-g81a45f2d8
git commit: 81a45f2
Reported by: nightly-eval via ailang messages