Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory a...

Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Role-Conditioned Sub-Token Routing (RoleSub) learns how to compress the value representations of retained tokens, outperforming a trained token-only control in 33 of 36 settings at matched visual-KV budgets.Source: arXiv cs.LGhttps://arxiv.org/abs/2608.18410#MachineLearning

Read Original

Related