Gorio Tech Blog search

VideoBERT - A Joint Model for Video and Language Representation Learning, CBT(Learning Video Representations using Contrastive Bidirectional Transformer) 논문 설명

|

이 글에서는 Google Research에서 발표한 VideoBERT(와 CBT) 논문을 간략하게 정리한다. VideoBERT: A Joint Model for Video and Language Representation Learning 논문 링크: VideoBERT: A Joint Model for Video and Language Representation Learning Github: maybe, https://github.com/ammesatyajit/VideoBERT 2019년 9월(Arxiv), ICCV 2019 Google Research Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and...

Comment  Read more

VL-BERT, ViL-BERT 논문 설명(VL-BERT - Pre-training of Generic Visual-Linguistic Representations, ViLBERT - Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks)

|

이 글에서는 VL-BERT와 ViLBERT 논문을 간략하게 정리한다. VL-BERT: Pre-training of Generic Visual-Linguistic Representations 논문 링크: VL-BERT: Pre-training of Generic Visual-Linguistic Representations Github: https://github.com/jackroos/VL-BERT 2020년 2월(Arxiv), ICLR 2020 University of Science and Technology of China, Microsoft Research Asia Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, Jifeng...

Comment  Read more

Bert4Rec(Sequential Recommendation with BERT) 설명

|

이번 글에서는 BERT 구조를 차용하여 추천 알고리즘을 구성해본 Bert4Rec이란 논문에 대해 다뤄보겠습니다. 논문 원본은 이 곳에서 확인할 수 있습니다. 본 글에서는 핵심적인 부분에 대해서만 살펴보겠습니다. Bert4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer 설명 1. Background RNN을 필두로 한 left-to-right unidirectional model은 user behavior sequence를 파악하기에 충분하지 않습니다. 왜냐하면...

Comment  Read more

metapath2vec(Scalable Representation Learning for Heterogeneous Networks) 설명

|

이번 글에서는 Heterogenous Network에서 node representation을 학습하는 metapath2vec이란 논문에 대해 다뤄보겠습니다. 논문 원본은 이 곳에서 확인할 수 있습니다. 본 글에서는 핵심적인 부분에 대해서만 살펴보겠습니다. 그리고 관련하여 학습/분석 코드는 이 곳에 작성해두었으니 참고하셔도 좋을 것 같습니다. metapath2vec: Scalable Representation Learning for Heterogeneous Networks 설명 1. Introduction word2vec 기반의 network representation learning...

Comment  Read more

ViT(Vision Transformer) 논문 설명(An Image is Worth 16x16 Words - Transformers for Image Recognition at Scale)

|

이 글에서는 ViT(Vision Transformer) 논문을 간략하게 정리한다. ViT(An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) 논문 링크: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Github: https://github.com/google-research/vision_transformer 2020년 10월, ICLR 2021 Google Research, Brain Team Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk...

Comment  Read more