A review of state of the art vision-language models such as CLIP, DALLE, ALIGN and SimVL
Vision Language models: towards multi-modal deep learning
Sergios Karagiannakos
Tags
A review of state of the art vision-language models such as CLIP, DALLE, ALIGN and SimVL