### Lecture 11: Multimodal I

1. `vit.ipynb`: Vision Transformer architecture 
2. `clip.ipynb`: using CLIP for image classification