AI Research Atlas
Preprint / conference paper

Multimodal models

CLIP connects images with natural language

Contrastive learning from image-text pairs enabled transfer through natural-language descriptions.

Alec Radford and colleagues

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

CLIP learned to associate images with their accompanying text. By comparing an image with text descriptions of candidate classes, the model could perform zero-shot classification without fitting a separate classifier to each evaluated dataset.

What this does not establish

Transfer performance depends on the task and prompts. Associations learned from web data are not equivalent to grounded or unbiased understanding.

Why this date?

The arXiv record lists a first submission of 26 February 2021, despite its 2103 identifier.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.