Computational linguistics

A search system misses matches between visually identical text strings because they use different canonical Unicode sequences. Which preprocessing choice most directly addresses this issue?