Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de...