Learning Neural Volumetric Pose Features for Camera Localization Paper • 2403.12800 • Published Mar 19, 2024 • 3
HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting Paper • 2412.03844 • Published Dec 5, 2024 • 3
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering Paper • 2510.14605 • Published Oct 16, 2025 • 8
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models Paper • 2509.17664 • Published Sep 22, 2025 • 4
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering Paper • 2602.23952 • Published Feb 27 • 6
HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization Paper • 2606.20097 • Published Jun 18 • 19
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop Paper • 2605.18746 • Published May 18 • 8
Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated Aug 11 • 200
Innovator-VL Collection A Multimodal Large Language Model for Scientific Discovery • 9 items • Updated Mar 5 • 6
SSR Collection Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning • 5 items • Updated Mar 2 • 2
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs Paper • 2512.16584 • Published Dec 18, 2025 • 4
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Paper • 2508.21113 • Published Aug 28, 2025 • 111