AI Research Atlas
Preprint / conference paper

Language

BERT learns bidirectional language representations

Masked pretraining helped a bidirectional Transformer transfer to many language tasks.

Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

BERT learned representations using context on both sides of masked tokens, then adapted those representations to downstream tasks through fine-tuning. The approach reduced the need for elaborate task-specific architectures in the evaluated settings.

What this does not establish

BERT is an encoder model with a different training setup from an autoregressive text generator. Benchmark gains do not establish unrestricted understanding.

Why this date?

The preprint appeared on 11 October 2018; the conference paper followed in 2019.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.