Computational linguistics

A model is trained to predict a deliberately hidden token using tokens both before and after the hidden position. Which description best distinguishes this setup from left-to-right autoregressive prediction?