Computational linguistics

Two language models are evaluated on the same text, but one tokenizer splits words into many more tokens. Why should their token-level perplexities not be compared without qualification?