Computational linguistics

A researcher reports next-token performance on examples that were also used to fit the model. Why is this not a reliable estimate of performance on unseen data?