Nightly eval REGRESSION: benchmark 'cli_args' passed ALL trials in the previous run but failed BOTH trials on 2026-07-18.
Error category: [logic_error]
Model: opencode-qwen3-5-35b-a3b-mxfp8 (local, tiers:smoke,core)
Results: /tmp/nightly_eval_20260718 (prev: /tmp/nightly_eval_20260717_rag_on/agent)
It was solid last night, so this is a fresh solid->broken regression — investigate.
Binary info (auto-attached):
ailang version: v0.29.2-364-g3b77bc036
git commit: 3b77bc0
Reported by: nightly-eval via ailang messages
Nightly eval REGRESSION: benchmark 'cli_args' passed ALL trials in the previous run but failed BOTH trials on 2026-07-18.
Error category: [logic_error]
Model: opencode-qwen3-5-35b-a3b-mxfp8 (local, tiers:smoke,core)
Results: /tmp/nightly_eval_20260718 (prev: /tmp/nightly_eval_20260717_rag_on/agent)
It was solid last night, so this is a fresh solid->broken regression — investigate.
Binary info (auto-attached):
ailang version: v0.29.2-364-g3b77bc036
git commit: 3b77bc0
Reported by: nightly-eval via ailang messages