10 Dec 2021
|
Transformer
Vision Transformer
이 글에서는 ViT(Vision Transformer) 논문을 간략하게 정리한다. ViT(An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) 논문 링크: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Github: https://github.com/google-research/vision_transformer 2020년 10월, ICLR 2021 Google Research, Brain Team Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk...
10 Dec 2021
|
Transformer
Vision Transformer
Video Transformers
이 글에서는 Transformer를 기반으로 Vision 문제를 푸는 모델인 ViViT(Video ViT: ViViT - A Video Vision Transformer), MTN, TimeSFormer, MViT 논문을 간략하게 정리한다. ViT(An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) 논문 링크: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Github: https://github.com/google-research/vision_transformer...
09 Dec 2021
|
Contrastive Learning
SimCLR
이 글에서는 Contrastive Learning을 간략하게 정리한다. Contrastive Learning 어떤 item들의 “차이”를 학습해서 그 rich representation을 학습하는 것을 말한다. 이 “차이”라는 것은 어떤 기준에 의해 정해진다. Contrastive Learning은 Positive pair와 Negative pair로 구성된다. 단, Metric Learning과는 다르게 한 번에 3개가 아닌 2개의 point를 사용한다. 한 가지 예시는, 같은 image에 서로 다른...
06 Dec 2021
|
Metric Learning
Loss Functions
이 글에서는 Metric Learning을 간략하게 정리한다. Metric Learning 한 문장으로 요약하면, Object간에 어떤 거리 함수를 학습하는 task 이다. 예를 들어 아래 이미지들을 보자. 뭔가 “이미지 간 거리”를 생각해보면, 1번째 이미지와 2번째 이미지는 거리가 가까울 것 같다. 이와는 대조적으로, 3번째 이미지는 다른 두 개의 이미지보다 거리가 멀 것 같다. 이런 관계를...
06 Dec 2021
|
Attention Mechanisms
Paper Review
Video Understanding
이 글에서는 Attention 기반 Video (Classification) Model을 간략히 소개한다. Multi-LSTM 논문 링크: Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos LRCN과 비슷하다. 다른 점은, Multiple Input: LSTM에 입력이 1개의 frame이 아니라 N개의 최근 frame에 대해 attention을 적용한다. Query: LSTM의 이전 hidden state $h_{i-1}$ Key=value: $N$개의 input frame...