Computational linguistics

Across many predictions assigned a probability near 0.8 to the correct next token, which finding would support calibration at that probability level?