Computational linguistics

A language-processing project needs both user-perceived character counts and word-level statistics. Which design best respects the difference between these goals?