NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI

NVIDIA today released Cosmos Reason 2, a state-of-the-art open reasoning vision-language model that enables robots and AI agents to see, understand, and act in the physical world with human-like reasoning.
NVIDIA today released Cosmos Reason 2, the latest advancement in open, reasoning vision language models for physical AI. Cosmos Reason 2 surpasses its previous version in accuracy and tops the Physical AI Bench and Physical Reasoning leaderboards as the #1 open model for visual understanding.
Since their introduction, vision-language models have rapidly improved at tasks like object and pattern recognition in images. But they still struggle with tasks humans find natural, like planning several steps ahead, dealing with uncertainty or adapting to new situations. Cosmos Reason is designed to close this gap by giving robots and AI agents stronger common sense and reasoning to solve complex problems step by step.
Cosmos Reason 2 is a state-of-the-art, open reasoning vision-language model (VLM) that enables robots and AI agents to see, understand, plan, and act in the physical world like humans. It uses common sense, physics, and prior knowledge to recognize how objects move across space and time to handle complex tasks, adapt to new situations, and figure out how to solve problems step by step.
Key features include improved spatio-temporal understanding, optimized performance with 2B and 8B parameter sizes, and support for 2D/3D point localization, bounding box coordinates, and OCR. It also features a massive 256K input token context window.
Major companies are already integrating these tools. Salesforce is using it for workplace safety via video analytics, while Uber is leveraging it for autonomous vehicle training data annotation, seeing double-digit improvements in evaluation metrics. Additionally, NVIDIA introduced Cosmos Predict for future-state forecasting and GR00T N1.6 for humanoid robot control.
Source: Hugging Face Blog

















