clip
Contrastive Language-Image Pre-training — 4억 (image, text) 대조 학습으로 공유 임베딩. 텍스트로 제로샷 분류 (OpenAI, ICML 2021).