Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text
Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves fr...