Autonomy-of-Heads proposes a data-free sparse attention method that reduces long-context LLM inference latency by up to ...

Autonomy-of-Heads proposes a data-free sparse attention method that reduces long-context LLM inference latency by up to 66% and KV-cache memory by 50% at 256K tokens, while retaining 96.5% of full attention performance at 50% sparsitySource: arXiv cs.CLhttps://arxiv.org/abs/2608.06849#MachineLearning

Read Original

Related