VisionGPT2
Combining ViT and GPT-2 for image captioning. Trained on MS-COCO. The model was implemented mostly from scratch.
// repository documentation
Was this content helpful?
(0 ratings)