31,430 frozen trials. 11 model identifiers. 4 providers. Models tested: gpt-4-0613, gpt-5.2-2025-12-11, gpt-5.5-2026-04-23, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, claude-opus-4-6, claude-fable-5, claude-opus-5, gemini-3.5-flash, kimi-k3 11,658 Voids. Strict matched pairs: 2,505/4,290 null arms produced Voids. 0/4,290 matched controls did. 9,093 were normal-stop Voids. At 16,000 tokens: 313/500 were still Voids. 0 were budget-stop Voids. “It’s just instruction following” is already considered in the paper. The question is simple: Does that explanation account for the full result? Matched asymmetry. Cross-provider behavior. Normal-stop zero-byte executions. High-token persistence. Ablations. Logical binding-condition contrasts. Separate refusal states. Scrutinize it. Reproduce it. Let’s discuss. submitted by /u/rayanpal_
Originally posted by u/rayanpal_ on r/ArtificialInteligence
