Computational linguistics

A language model gives token A the highest probability at the first step. After A, all continuations score poorly; a lower-probability first token B has a much stronger continuation. What limitation of greedy decoding does this illustrate?