An ML search loop that can't overfit its own eval — the gate is in code, not a prompt
I kept seeing agent and eval demos where the honesty — held-out discipline, "no metric gaming" — lives in a prompt, or in a paper's methodology section. So I tried to build the opposite: a search loop
Jul 25, 20263 min read1
