Disentangling the Expressivity of RoPE
Selim Jerad ⋅ Anej Svete ⋅ Jiaoda Li ⋅ Ryan Cotterell
Abstract
Rotary position embeddings (RoPE) inject positional information in the attention mechanism by rotating the keys and queries based on their positions. RoPE has become the go-to positional encoding scheme due to its strong empirical performance, but little is known theoretically about how it affects transformers' expressivity---what they can and cannot compute. We provide a new perspective on RoPE by comparing the expressivity of RoPE transformers to those with no positional encodings (NoPE). We build on prior work that describes NoPE transformers with a fragment of first-order logic, and precisely characterize soft-attention, finite-precision RoPE transformers: They correspond to the same fragment of first-order logic augmented with modular predicates. We also prove that, up to a precision-dependent length bound, RoPE transformers can simulate the $\texttt{yesterday}$ temporal operator. Altogether, our findings shed light on the exact capabilities the most used position encoding scheme buys, bringing theoretical transformer expressivity characterizations closer to models used in practice.
Successful Page Load