Hello, everyone. Attention becomes expensive very quickly as more text is given to an AI model. Can a Mac GPU make it faster when every token is restricted to looking only at nearby tokens?

Source: [Dev.to](https://dev.to/kiarina/testing-pytorch-213-mps-flexattention-on-m1-max-up-to-783x-faster-for-sparse-attention-1p7d)

Sponsored