A held-out test from 2021 through 2025 is the useful part of this result. IC is the correlation between model scores and the returns that follow. The paper reports per-stock raw IC of +0.0843, compared with +0.0613 for its strongest GRU baseline. In AQuA, the claim is not that a language model traded the market. It is that an agent-guided research process found a stronger model configuration under a fixed evaluation setup. There is an obvious boundary. The result is simulated and has no live-trading validation. A held-out lift can show that the search protocol found something better under this evaluator, but it cannot establish that the result is tradable. For an agent-generated model result, which evidence should carry the most weight: an untouched test period, an independent reproduction, released implementation details, or forward performance? submitted by /u/AcanthisittaOk1699
Originally posted by u/AcanthisittaOk1699 on r/ArtificialInteligence
