CLIP

E95184

CLIP is an OpenAI model that learns joint representations of images and text, enabling tasks like zero-shot image classification and natural language-based image retrieval.

AI illustration

How this image was made

AI-generated illustration of CLIP

This AI-generated illustration was produced by black-forest-labs/FLUX.2-dev (1024x1024) from a prompt written by openai/gpt-oss-120b from the entity's label + description.

Prompt

Generate an image of CLIP (CLIP is an OpenAI model that learns joint representations of images and text, enabling tasks like zero-shot image classification and natural language-based image retrieval.)

All labels observed (7)

How this entity was disambiguated

Statements (56)

Predicate Object
instanceOf contrastive learning model ⓘ
multimodal machine learning model ⓘ
vision-language model ⓘ
architectureComponent image encoder ⓘ
text encoder ⓘ
capability cross-modal retrieval ⓘ
learning from natural language supervision ⓘ
open-vocabulary recognition ⓘ
prompt-based classification ⓘ
zero-shot transfer to downstream vision tasks ⓘ
developer OpenAI ⓘ
fullName Contrastive Language–Image Pre-training ⓘ
linked to: CLIP
imageEncoderType ResNet ⓘ
Vision Transformer ⓘ
linked to: ViT
input image ⓘ
natural language text prompt ⓘ
inspired subsequent vision-language models ⓘ
introducedBy Aditya Ramesh ⓘ
Alec Radford ⓘ
Amanda Askell ⓘ
Chris Hallacy ⓘ
Gabriel Goh ⓘ
Girish Sastry ⓘ
Gretchen Krueger ⓘ
Ilya Sutskever ⓘ
Jack Clark ⓘ
Jong Wook Kim ⓘ
Pamela Mishkin ⓘ
Sandhini Agarwal ⓘ
learningParadigm contrastive learning ⓘ
self-supervised learning ⓘ
license OpenAI model license ⓘ
lossFunction InfoNCE-style loss ⓘ
contrastive loss ⓘ
modality image ⓘ
text ⓘ
organization OpenAI ⓘ
output joint embedding vectors for images and text ⓘ
similarity scores between images and text ⓘ
pretrainingDataType image-text pairs ⓘ
property aligns image and text embeddings in a shared space ⓘ
does not require task-specific fine-tuning for many tasks ⓘ
uses cosine similarity in embedding space ⓘ
publicationTitle Learning Transferable Visual Models From Natural Language Supervision ⓘ
linked to: CLIP
publicationType arXiv preprint ⓘ
publicationYear 2021 ⓘ
task image representation learning ⓘ
image-text matching ⓘ
natural language-based image retrieval ⓘ
text representation learning ⓘ
zero-shot image classification ⓘ
textEncoderType Transformer ⓘ
trainingObjective maximize similarity of matching image-text pairs ⓘ
minimize similarity of non-matching image-text pairs ⓘ
usedFor as a backbone in multimodal systems ⓘ
downstream fine-tuning for vision tasks ⓘ

How these facts were elicited

Referenced by (15)

Full triples — surface form annotated when it differs from this entity's canonical label.

DALL·E → relatedTo → CLIP ⓘ
CLIP → fullName → Contrastive Language–Image Pre-training ⓘ
linked to: CLIP
CLIP → publicationTitle → Learning Transferable Visual Models From Natural Language Supervision ⓘ
linked to: CLIP
Jong Wook Kim → knownFor → CLIP ⓘ
Jong Wook Kim → contributedTo → CLIP: Connecting Text and Images ⓘ
linked to: CLIP
Aditya Ramesh → associatedWith → OpenAI CLIP ⓘ
linked to: CLIP
Gabriel Goh → notableWork → CLIP ⓘ
Gabriel Goh → coAuthorOf → “Learning Transferable Visual Models From Natural Language Supervision” ⓘ
linked to: CLIP
Gabriel Goh → developedAt → CLIP, at OpenAI ⓘ
linked to: CLIP
“Learning Transferable Visual Models From Natural Language Supervision” → mainSubject → CLIP ⓘ
subject linked to: Gabriel Goh