2018.05.01
NAIST ⾃自然⾔言語処理理学研究室
D2 Masayoshi Kondo
論論⽂文紹介-‐‑‒ About Neural Summarization@2018
Attention Is All You Need
NIPSʼ’17
Ashish Vaswani
Google Brain
Niki Parmar
Google Research
Jakob Uszkoreit
Google Research
Llion Jones
Google Research
Aidan N. Gomez
University of Toronto
Łukasz Kaiser
Google Brain
Illia Polosukhin
Noam Shazeer
Google Brain
k
• Self-‐‑‒attention層 と Recurrent層 or Convolution層 の⽐比較を⾏行行う.
• 可変⻑⾧長系列列(x1, x2, …, xn) から 同じ⻑⾧長さの系列列(z1, z2,…, zn) へ
の変換を考える (x,z ∈ Rd).
• Self-‐‑‒attentionを⽤用いることの強みは、3つの必要性を考えれば分かる.
13 : Why Self-‐‑‒Attention
この章では:
1. ⼀一層あたりの総合計算量量
( total computational complexity per layer )
2. 要求される最⼩小の系列列演算数から測られる並列列可能な計算量量
( the amount of computation that can be parallelized, as measured by
the minimum number of sequential operations required. )
3. ネットワーク内部の⻑⾧長距離離依存関係間のパスの⻑⾧長さ
( the path length between long-‐‑‒range dependencies
in the network )
22.
14 : Why Self-‐‑‒Attention
多くの系列列変換タスクで、⻑⾧長距離離の依存性を学習することは重要な課題.
3. ネットワーク内部の⻑⾧長距離離依存関係間のパスの⻑⾧長さ
( the path length between long-‐‑‒range dependencies
in the network )
依存関係性を学習する能⼒力力に影響を与える重要な因⼦子:
ネットワーク内の前⽅方向・後⽅方向に信号が伝搬しなければならない距離離の⻑⾧長さ
⼊入出⼒力力系列列における任意の位置の組み合わせ間のパス(経路路)が短くなれば
なるほど、⻑⾧長距離離依存性の学習が容易易になる.
-‐‑‒ Gradient flow in recurrent nets: the difficulty of learning long-‐‑‒term dependencies.
異異なる層で構成されるネットワーク内部の2つの⼊入出⼒力力の位置の間の
最⼤大経路路⻑⾧長を⽐比較.
【考察・検証】