Computational linguistics

A text collection has no spaces between many adjacent words, but its language has established word-boundary conventions. What is the best general approach for creating word tokens?