Gorio Tech Blog search

ViT(Vision Transformer) 논문 설명(An Image is Worth 16x16 Words - Transformers for Image Recognition at Scale)

|

이 글에서는 ViT(Vision Transformer) 논문을 간략하게 정리한다. ViT(An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) 논문 링크: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Github: https://github.com/google-research/vision_transformer 2020년 10월, ICLR 2021 Google Research, Brain Team Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk...

Comment  Read more

ViViT(Video ViT, ViViT - A Video Vision Transformer), MTN, TimeSFormer, MViT 논문 설명

|

이 글에서는 Transformer를 기반으로 Vision 문제를 푸는 모델인 ViViT(Video ViT: ViViT - A Video Vision Transformer), MTN, TimeSFormer, MViT 논문을 간략하게 정리한다. ViT(An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) 논문 링크: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Github: https://github.com/google-research/vision_transformer...

Comment  Read more

Contrastive Learning, SimCLR 논문 설명(SimCLRv1, SimCLRv2)

|

이 글에서는 Contrastive Learning을 간략하게 정리한다. Contrastive Learning 어떤 item들의 “차이”를 학습해서 그 rich representation을 학습하는 것을 말한다. 이 “차이”라는 것은 어떤 기준에 의해 정해진다. Contrastive Learning은 Positive pair와 Negative pair로 구성된다. 단, Metric Learning과는 다르게 한 번에 3개가 아닌 2개의 point를 사용한다. 한 가지 예시는, 같은 image에 서로 다른...

Comment  Read more

Metric Learning 설명

|

이 글에서는 Metric Learning을 간략하게 정리한다. Metric Learning 한 문장으로 요약하면, Object간에 어떤 거리 함수를 학습하는 task 이다. 예를 들어 아래 이미지들을 보자. 뭔가 “이미지 간 거리”를 생각해보면, 1번째 이미지와 2번째 이미지는 거리가 가까울 것 같다. 이와는 대조적으로, 3번째 이미지는 다른 두 개의 이미지보다 거리가 멀 것 같다. 이런 관계를...

Comment  Read more

Attention based Video Models

|

이 글에서는 Attention 기반 Video (Classification) Model을 간략히 소개한다. Multi-LSTM 논문 링크: Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos LRCN과 비슷하다. 다른 점은, Multiple Input: LSTM에 입력이 1개의 frame이 아니라 N개의 최근 frame에 대해 attention을 적용한다. Query: LSTM의 이전 hidden state $h_{i-1}$ Key=value: $N$개의 input frame...

Comment  Read more