Adapting the Transformer: Shift from machine translation to generalized representation with BERT
By leveraging the multi-head attention layer from the encoder only, BERT introduced the CLS token to capture sequence-wide context, facilitating downstream tasks like sentiment extraction without full model retraining.
This architecture strategically utilizes the multi-head attention layer from the encoder to enable self-attention across all input tokens, allowing them to attend to one another bidirectionally.
BERT also introduced the CLS (classification) token, a special input token designed to capture aggregated contextual information from the entire input sequence, which proved crucial for various downstream tasks like sentiment extraction without requiring full model retraining.


