Block-Recurrent Transformers (arXiv)

“We introduce the Block-Recurrent Transformer, which applies a transformer layer in a recurrent fashion along a sequence, and has linear complexity with respect to sequence length. Our recurrent cell operates on blocks of tokens rather than single tokens, and leverages parallel computation within a block in order to make efficient use of accelerator hardware. The cell itself is strikingly simple. It is merely a transformer layer: it uses self-attention and cross-attention to efficiently compute a recurrent function over a large set of state vectors and tokens. Our design was inspired in part by LSTM cells, and it uses LSTM-style gates, but it scales the typical LSTM cell up by several orders of magnitude.
Our implementation of recurrence has the same cost in both computation time and parameter count as a conventional transformer layer, but offers dramatically improved perplexity in language modeling tasks over very long sequences. Our model out-performs a long-range Transformer XL baseline by a wide margin, while running twice as fast. We demonstrate its effectiveness on PG19 (books), arXiv papers, and GitHub source code.”

https://arxiv.org/abs/2203.07852

Block-Recurrent Transformers (arXiv)

Byauthor

Like this:

By author

Related Post

Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages (arXiv)

Social Hatred: Efficient Multimodal Detection of Hatemongers (arXiv)

The Role of Context in Detecting the Target of Hate Speech (ACL Anthology)

Leave a Reply Cancel reply

LATEST NEWS

Criminalising Hate Speech: A Comparative Study (SSRN)

Counterspeech encouraging users to adopt the perspective of minority groups reduces hate speech and its amplification on social media (scientific reports)