Collections
Discover the best community collections!
Collections including paper arxiv:2607.21553
-
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Paper • 2604.12012 • Published • 16 -
InferenceSupport
💥649Discussions about the Inference Providers feature on the Hub
-
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
Paper • 2607.20709 • Published • 35 -
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
Paper • 2607.21553 • Published • 39
-
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Paper • 2605.23902 • Published • 47 -
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
Paper • 2605.30263 • Published • 60 -
From Pixels to Words -- Towards Native One-Vision Models at Scale
Paper • 2605.28820 • Published • 76 -
open-thoughts/AgentTrove
Viewer • Updated • 1.7M • 2.24k • 195
-
AI for Auto-Research: Roadmap & User Guide
Paper • 2605.18661 • Published • 72 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 15 -
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper • 2605.16865 • Published • 10 -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper • 2603.28069 • Published • 9
-
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Paper • 2605.23902 • Published • 47 -
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
Paper • 2605.30263 • Published • 60 -
From Pixels to Words -- Towards Native One-Vision Models at Scale
Paper • 2605.28820 • Published • 76 -
open-thoughts/AgentTrove
Viewer • Updated • 1.7M • 2.24k • 195
-
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Paper • 2604.12012 • Published • 16 -
InferenceSupport
💥649Discussions about the Inference Providers feature on the Hub
-
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
Paper • 2607.20709 • Published • 35 -
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
Paper • 2607.21553 • Published • 39
-
AI for Auto-Research: Roadmap & User Guide
Paper • 2605.18661 • Published • 72 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 15 -
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper • 2605.16865 • Published • 10 -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper • 2603.28069 • Published • 9