Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models
As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). M...