ATPA: A holdout must resemble the intended prediction problem
A random holdout measures a different task from predicting into a later period when trends or process changes matter. Repeated customer records can also leak entity-specific information across training and validation sets.
Worked example or practice scenario
If the goal is next-year claims prediction, train on earlier periods and validate on a later period as an original diagnostic exercise. If the goal involves entirely new customers, ensure the split does not place the same customer in both sets. The appropriate design depends on deployment.
Try this next
Describe what the validation split represents and what it cannot measure. Compare performance by period and segment. Fit preprocessing using the training data only, and keep a final assessment set outside iterative tuning decisions.
Reading sources
ActNet editorial guide · October 1, 2026 · Original illustrative examples.