Archived writing · retained for reference.
How I evaluate a market hypothesis
A market hypothesis starts with a mechanism: which observation changes, what action becomes possible, and why should the resulting payoff differ from a baseline? I write down those assumptions before interpreting a result.
Define the measurement
A decision moment, a filled order and a market window are different units. A useful result names its unit, includes the full eligible sample and states what information was available at the decision. Data gaps and arrival clocks belong in that definition.
Test alternatives
Chronological evaluation checks behavior on later data. Matched placebos test whether the proposed information matters beyond a generic action such as withdrawing orders. Clustered uncertainty keeps related observations together. Stability checks ask whether one window or parameter choice drives the conclusion. Each result needs its own checks; a protocol is not proof that every study passed.
Check what the model represents
A replay can model activation and queue depletion while missing the response to our own orders. A positive estimate remains dependent on those assumptions. A broken paper run can reveal an implementation failure without establishing the economics of the intended strategy.
Publish the revision
When an assumption fails, the earlier conclusion needs an explicit revision. In the crypto study, corrected maker activation and fresh data changed the static-front policy decision. The public repository records both the experiment's scope and what its aggregate verifier can reproduce.