When strategy rules are created, the data you have already seen influences both the idea and the parameter choices. Reusing that same data as the final test is like reading the exam, choosing the answer, and then grading yourself on the same questions.

In-sample is for building; out-of-sample is for testing.

In-sample data is where ideas are explored and structured. Final performance should then be evaluated on a separate out-of-sample period the strategy has not seen. That helps distinguish a rule that memorized noise from one that carries into new data.

TIME ORDER
Past data→Design / tune→Freeze rules→Future OOS→Evaluate

A weak OOS result should not be hidden

Strong in-sample performance and weak OOS performance is common. If the OOS period is then excluded or the rule is changed retroactively, the test set turns back into training data. If changes are needed, define a new version and a new test period instead.

Keeping time order is key

Random splits are not always wrong, but time-series strategies require especially strict time boundaries so that future information does not leak into earlier training. Changes in market structure and data structure should also be recorded.