VisionTransformer

(★ 101)

A complete easy to follow implementation of Google's Vision Transformer proposed in "AN IMAGE IS WORTH 16X16 WORDS". This pytorch implementation has comments for better understanding.

  • Google_ViT.py
  • README.md
  • ViT.png
// repository documentation