NVIDIA
NVIDIA Agent Skills v1.0
NVIDIA Agent Skills are a collection of official, NVIDIA-verified skills designed to enhance AI agents' capabilities by integrating them with NVIDIA's CUDA-X libraries, AI Blueprints, and platform tools. These skills enable AI agents to perform tasks such as optimization, data processing, and simulation, thereby accelerating the development of AI applications across various domains.
VibeTensor v1.0
VibeTensor is an open-source deep learning system software stack developed by NVIDIA, generated entirely by large language model (LLM)-powered coding agents under high-level human guidance. It offers a comprehensive solution for deep learning tasks, integrating various components from tensor libraries to CUDA memory management.
PersonaPlex-7B-v1 v7B-v1
PersonaPlex-7B-v1 is a real-time speech-to-speech conversational model developed by NVIDIA, designed to enable natural and full-duplex conversations. It operates on continuous audio, performing both speech understanding and generation simultaneously, allowing for dynamic interactions such as interruptions and rapid turn-taking.
Dynamo v0.9.0 vv0.9.0
Dynamo v0.9.0 is an open-source inference library developed by NVIDIA, designed to accelerate and scale AI reasoning models within AI factories. This release marks a significant infrastructure overhaul, introducing features like FlashIndexer and enhanced support for multimodal models, while removing dependencies on NATS and etcd.
Qwen3-8B-DMS-8x vQwen3-8B-DMS-8x
Qwen3-8B-DMS-8x is a derivative of Qwen3-8B that integrates Dynamic Memory Sparsification (DMS) with an 8x compression ratio during inference. DMS adaptively sparsifies the key-value (KV) cache to reduce memory footprint and improve throughput and latency for long-context and reasoning generations. The method learns per-head eviction policies that interpolate between a sliding window over the last 512 tokens and full attention. Inference-time code is provided with the checkpoint.
DeepStream SDK v9.1
NVIDIA DeepStream SDK is a powerful streaming analytics toolkit designed for AI-based video and image understanding. It provides a GStreamer-based framework that enables developers to build highly efficient, multi-stream, and multi-model inference pipelines on NVIDIA dGPU and Jetson platforms.
NeMo Speech v2.7.3
NVIDIA NeMo Speech is a scalable generative AI framework specifically engineered for researchers and developers working on Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and speech-enabled LLMs. It enables users to efficiently build, customize, and deploy state-of-the-art speech AI models by utilizing pre-trained checkpoints and modular code components.
NitroGen v1.0
NitroGen is a unified vision-to-action foundation model developed by NVIDIA, designed to play video games directly from raw frames. Trained on 40,000 hours of gameplay across over 1,000 games, it maps RGB video footage to gamepad actions, enabling generalist gaming agents to perform tasks in various gaming environments. ([huggingface.co](https://huggingface.co/nvidia/NitroGen?utm_source=openai))
NVIDIA BlueField-4-Powered Inference Context Memory Storage Platform
NVIDIA's BlueField-4-Powered Inference Context Memory Storage Platform is an AI-native storage infrastructure designed to accelerate and scale agentic AI workloads. It extends GPU memory capacity and enables high-speed sharing of context data across AI systems, improving performance and power efficiency. ([developer.nvidia.com](https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/?utm_source=openai))
CUDA 13.2 v13.2
CUDA 13.2 is the latest release of NVIDIA's parallel computing platform and application programming interface (API) model, designed to leverage the power of NVIDIA GPUs for high-performance computing tasks. This version introduces significant enhancements, including full support for CUDA Tile on various GPU architectures and advanced features in cuTile Python, aiming to simplify GPU programming and improve developer productivity.
Nemotron Speech ASR v0.6b
NVIDIA has released Nemotron Speech ASR, a streaming English transcription model designed for low-latency voice agents and live captioning. This 600 million parameter model utilizes a cache-aware FastConformer encoder and an RNNT decoder, optimized for both streaming and batch workloads on modern NVIDIA GPUs.
Nemotron 3.5 Lightning v3.5
NVIDIA Nemotron 3.5 Lightning 30B-A3B-NVFP4 is a high-performance, latency-optimized large language model featuring a hybrid Mamba-2 and Mixture-of-Experts (MoE) architecture. Designed for efficient agentic workflows, it utilizes 30B total parameters with 3B active parameters per token to deliver fast, accurate execution for specialized tasks.
NVIDIA Ising
NVIDIA Ising is the world's first family of open-source AI models for quantum computing, designed to accelerate quantum processor calibration and quantum error correction. Named after the landmark Ising mathematical model, it delivers up to 2.5× faster and 3× more accurate quantum error correction decoding than the current open-source standard (pyMatching), and automates continuous processor calibration from days to hours.
NVIDIA Earth-2 v1.0
NVIDIA Earth-2 is a comprehensive suite of open-source models and tools designed to revolutionize weather and climate forecasting through accelerated AI technologies. It offers a fully open, accelerated weather AI software stack, enabling organizations worldwide to develop and deploy their own forecasting systems.
NVIDIA Nemotron 3 Nano Omni v3 Nano Omni
NVIDIA's new high-efficiency multimodal model designed for edge AI agents. Unifies vision, audio, and language into a single small-parameter model, delivering up to 9x better efficiency for real-time agentic workflows.
NVIDIA Cosmos 3
NVIDIA Cosmos 3 is a frontier open-source physical AI foundation model that unifies physical reasoning, world generation, and action generation in a single model using a Mixture-of-Transformers (MoT) architecture. It combines a Reasoner tower (VLM for understanding) and a Generator tower (diffusion-based video/action output), eliminating the need for multiple separate models. Available in two sizes — Cosmos 3 Nano (8B) and Cosmos 3 Super (32B) — with fully open model weights, training scripts, deployment tools, and six synthetic datasets on Hugging Face.
NemoClaw for LangChain Deep Agents Code v0.0.87
NemoClaw for LangChain Deep Agents Code is a governed blueprint that allows developers to run open-source coding agents using NVIDIA Nemotron 3 Ultra models. It provides a secure, audit-friendly environment for performing complex software engineering tasks like refactoring, dependency upgrades, and automated code maintenance.
Nemotron-RL Agentic Terminal Pivot v1
Nemotron-RL Agentic Terminal Pivot v1 is an open-source reinforcement learning dataset designed to train large language models for agentic command-line interface (CLI) tasks. It provides high-quality, expert-derived trajectories that enable models to perform complex software engineering operations, tool use, and reasoning within Linux environments.
Nemotron 3 Nano v3 Nano
Nemotron 3 Nano is an open-source language model designed for agentic AI applications, featuring a Mixture of Experts hybrid Mamba Transformer architecture with approximately 31.6 billion parameters. It offers efficient long-context reasoning capabilities, supporting up to 1 million tokens, and is optimized for multi-agent systems operating on extensive documents and codebases.
NeMo Agent Toolkit v1.4
NVIDIA NeMo™ Agent Toolkit is an open-source AI framework designed to build, profile, and optimize AI agents and tools across various frameworks. It enables unified, cross-framework integration for connected AI agent systems, helping enterprises efficiently scale agentic systems while maintaining reliability. ([developer.nvidia.com](https://developer.nvidia.com/agent-intelligence-toolkit?utm_source=openai))
NVIDIA Nemotron 3 Ultra v3 Ultra
NVIDIA Nemotron 3 Ultra is a 550-billion-parameter Mixture-of-Experts model designed to enhance the efficiency and speed of long-running agents in complex workflows. It combines advanced reasoning capabilities with high throughput, enabling agents to perform tasks faster and more cost-effectively.