Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
Abstract
Credit-addressable reasoning via executable code traces and localized reinforcement learning improves multimodal geometry reasoning by aligning credit assignment with structured reasoning events.
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-level reinforcement learning distributes a single terminal signal across the entire response. We introduce credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit. We instantiate this principle with Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events, and CE-GRPO, which selects event boundaries using structural priors and type-normalized entropy, samples complete continuations from shared prefixes, and converts outcome differences into localized advantages. Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by 8.09 and 3.43 points, respectively. Its relative advantage increases with the number of intermediate events, demonstrating the value of representation--optimization co-design for long, dependency-heavy multimodal reasoning.
Community
This work introduces credit-addressable reasoning for multimodal geometry, aiming to localize learning signals to the reasoning steps that actually change outcomes. We propose Code-CoT for structured, executable reasoning and CE-GRPO for event-level credit assignment, achieving 76.04% average accuracy across nine geometry benchmarks and outperforming trajectory-level GRPO by 3.43 points.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning (2026)
- AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning (2026)
- GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning (2026)
- Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering (2026)
- Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception (2026)
- MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems (2026)
- OPLD: On-Policy Latent Distillation for Multimodal Reasoning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.30457 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper