NVIDIA

2026NVIDIA

NVIDIA Agent Skills v1.0

Open Source
AI agents
NVIDIA

NVIDIA Agent Skills are a collection of official, NVIDIA-verified skills designed to enhance AI agents' capabilities by integrating them with NVIDIA's CUDA-X libraries, AI Blueprints, and platform tools. These skills enable AI agents to perform tasks such as optimization, data processing, and simulation, thereby accelerating the development of AI applications across various domains.

Integration with NVIDIA CUDA-X libraries
Access to AI Blueprints and platform tools
Support for tasks like optimization, data processing, and simulation
PricingFree
Version1.0
2026NVIDIA

VibeTensor v1.0

Open Source
Coding
Agentic AI

VibeTensor is an open-source deep learning system software stack developed by NVIDIA, generated entirely by large language model (LLM)-powered coding agents under high-level human guidance. It offers a comprehensive solution for deep learning tasks, integrating various components from tensor libraries to CUDA memory management.

PyTorch-style eager tensor library with a C++20 core supporting CPU and CUDA operations
Python interface via nanobind, and an experimental Node.js/TypeScript interface
Integrated tensor and storage system with schema-lite dispatcher
PricingFree
Version1.0
2026NVIDIA

PersonaPlex-7B-v1 v7B-v1

Open Source
AI Agents
Agentic AI

PersonaPlex-7B-v1 is a real-time speech-to-speech conversational model developed by NVIDIA, designed to enable natural and full-duplex conversations. It operates on continuous audio, performing both speech understanding and generation simultaneously, allowing for dynamic interactions such as interruptions and rapid turn-taking.

Full-duplex conversational capabilities
Customizable voice and role through prompts
Operates on continuous audio for real-time interaction
PricingFree
Version7B-v1
2026NVIDIA

Dynamo v0.9.0 vv0.9.0

Open Source
Infrastructure

Dynamo v0.9.0 is an open-source inference library developed by NVIDIA, designed to accelerate and scale AI reasoning models within AI factories. This release marks a significant infrastructure overhaul, introducing features like FlashIndexer and enhanced support for multimodal models, while removing dependencies on NATS and etcd.

FlashIndexer integration for improved indexing performance
Expanded support for multimodal models across all backends
Decoupling of infrastructure components for enhanced scalability
PricingFree
Versionv0.9.0
2026NVIDIA

Qwen3-8B-DMS-8x vQwen3-8B-DMS-8x

Open Source
ML

Qwen3-8B-DMS-8x is a derivative of Qwen3-8B that integrates Dynamic Memory Sparsification (DMS) with an 8x compression ratio during inference. DMS adaptively sparsifies the key-value (KV) cache to reduce memory footprint and improve throughput and latency for long-context and reasoning generations. The method learns per-head eviction policies that interpolate between a sliding window over the last 512 tokens and full attention. Inference-time code is provided with the checkpoint.

Integrates Dynamic Memory Sparsification (DMS) with 8x compression ratio during inference
Reduces memory footprint and improves throughput and latency for long-context and reasoning generations
Provides inference-time code with the checkpoint
VersionQwen3-8B-DMS-8x
RegionUnited Kingdom
2026NVIDIA

DeepStream SDK v9.1

Free
Computer Vision
AI Inference

NVIDIA DeepStream SDK is a powerful streaming analytics toolkit designed for AI-based video and image understanding. It provides a GStreamer-based framework that enables developers to build highly efficient, multi-stream, and multi-model inference pipelines on NVIDIA dGPU and Jetson platforms.

Hardware-accelerated video decoding, encoding, and inference
GStreamer-based framework for multi-stream/multi-model pipelines
Deep integration with TensorRT and TAO-trained models
Version9.1
RegionUnited States
2026NVIDIA

NeMo Speech v2.7.3

Open Source
Speech AI
Generative AI

NVIDIA NeMo Speech is a scalable generative AI framework specifically engineered for researchers and developers working on Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and speech-enabled LLMs. It enables users to efficiently build, customize, and deploy state-of-the-art speech AI models by utilizing pre-trained checkpoints and modular code components.

Supports Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) pipelines
Integrates seamlessly with multimodal Large Language Models (LLMs)
Offers low-latency, full-duplex voice interaction capabilities
PricingOpen source (Apache License 2.0)
Version2.7.3
2025NVIDIA

NitroGen v1.0

Open Source
AI Agents
RL

NitroGen is a unified vision-to-action foundation model developed by NVIDIA, designed to play video games directly from raw frames. Trained on 40,000 hours of gameplay across over 1,000 games, it maps RGB video footage to gamepad actions, enabling generalist gaming agents to perform tasks in various gaming environments. ([huggingface.co](https://huggingface.co/nvidia/NitroGen?utm_source=openai))

Trained on 40,000 hours of gameplay across over 1,000 games
Maps RGB video footage to gamepad actions
Supports a wide range of game genres
PricingFree
Version1.0
2026NVIDIA

NVIDIA BlueField-4-Powered Inference Context Memory Storage Platform

Paid
Infrastructure

NVIDIA's BlueField-4-Powered Inference Context Memory Storage Platform is an AI-native storage infrastructure designed to accelerate and scale agentic AI workloads. It extends GPU memory capacity and enables high-speed sharing of context data across AI systems, improving performance and power efficiency. ([developer.nvidia.com](https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/?utm_source=openai))

Extends GPU memory capacity for AI workloads
Enables high-speed sharing of context data across AI systems
Improves performance and power efficiency
RegionUnited States
2026NVIDIA

CUDA 13.2 v13.2

Open Source
Infrastructure
Agentic AI

CUDA 13.2 is the latest release of NVIDIA's parallel computing platform and application programming interface (API) model, designed to leverage the power of NVIDIA GPUs for high-performance computing tasks. This version introduces significant enhancements, including full support for CUDA Tile on various GPU architectures and advanced features in cuTile Python, aiming to simplify GPU programming and improve developer productivity.

Full support for CUDA Tile on compute capability 8.X (Ampere, Ada), 10.X, and 12.X (Blackwell) architectures
Enhanced cuTile Python with recursive functions, closures, custom reductions, type-annotated assignments, and improved array slicing
New `cudaMemcpyWithAttributesAsync` and `cudaMemcpy3DWithAttributesAsync` APIs for flexible memory transfers
PricingFree
Version13.2
2026NVIDIA

Nemotron Speech ASR v0.6b

Open Source
AI Agents
Agentic AI

NVIDIA has released Nemotron Speech ASR, a streaming English transcription model designed for low-latency voice agents and live captioning. This 600 million parameter model utilizes a cache-aware FastConformer encoder and an RNNT decoder, optimized for both streaming and batch workloads on modern NVIDIA GPUs.

Cache-aware FastConformer encoder with 24 layers
RNNT decoder for efficient transcription
Configurable context sizes for latency control
PricingFree
Version0.6b
2026NVIDIA

Nemotron 3.5 Lightning v3.5

Open Source
text-generation
mixture-of-experts

NVIDIA Nemotron 3.5 Lightning 30B-A3B-NVFP4 is a high-performance, latency-optimized large language model featuring a hybrid Mamba-2 and Mixture-of-Experts (MoE) architecture. Designed for efficient agentic workflows, it utilizes 30B total parameters with 3B active parameters per token to deliver fast, accurate execution for specialized tasks.

Hybrid Mamba-2 and Mixture-of-Experts (MoE) architecture
30B total parameters with 3B active parameters per token
1 million token context window
PricingFree (Open Weights)
Version3.5
2026NVIDIA

NVIDIA Ising

Paid
Quantum Computing
NVIDIA

NVIDIA Ising is the world's first family of open-source AI models for quantum computing, designed to accelerate quantum processor calibration and quantum error correction. Named after the landmark Ising mathematical model, it delivers up to 2.5× faster and 3× more accurate quantum error correction decoding than the current open-source standard (pyMatching), and automates continuous processor calibration from days to hours.

Ising Calibration — a vision language model that interprets measurements from quantum processors in real time, enabling AI agents to automate continuous calibration and reducing time needed from days to hours; adopted by IonQ, Atom Computing, Harvard, Fermi National Lab, and more
Ising Decoding — two 3D CNN model variants (speed-optimized and accuracy-optimized) for real-time quantum error correction; up to 2.5× faster and 3× more accurate than pyMatching, the current open-source industry standard
Fully open and customizable — models, tools, and training data available on GitHub, Hugging Face, and build.nvidia.com; integrates with NVIDIA CUDA-Q software platform and NVQLink QPU-GPU hardware interconnect; fine-tunable for specific hardware architectures
PricingFree (Open Source)
RegionGlobal
2026NVIDIA

NVIDIA Earth-2 v1.0

Open Source
LLMs
ML

NVIDIA Earth-2 is a comprehensive suite of open-source models and tools designed to revolutionize weather and climate forecasting through accelerated AI technologies. It offers a fully open, accelerated weather AI software stack, enabling organizations worldwide to develop and deploy their own forecasting systems.

Open-source models and tools for weather and climate AI
Accelerated AI technologies for faster and more accurate forecasting
Comprehensive suite including models like Earth-2 Medium Range, Nowcasting, and Global Data Assimilation
PricingFree
Version1.0
2026NVIDIA

NVIDIA Nemotron 3 Nano Omni v3 Nano Omni

Paid
NVIDIA
Multimodal

NVIDIA's new high-efficiency multimodal model designed for edge AI agents. Unifies vision, audio, and language into a single small-parameter model, delivering up to 9x better efficiency for real-time agentic workflows.

Unifies vision, audio, and language in one small-parameter model
Up to 9x better efficiency vs fragmented multi-model stacks for edge agents
Built on NVIDIA Parakeet speech encoder — designed for real-time agentic use
PricingFree (open model)
Version3 Nano Omni
2026NVIDIA

NVIDIA Cosmos 3

Paid
Featured
NVIDIA
Physical AI

NVIDIA Cosmos 3 is a frontier open-source physical AI foundation model that unifies physical reasoning, world generation, and action generation in a single model using a Mixture-of-Transformers (MoT) architecture. It combines a Reasoner tower (VLM for understanding) and a Generator tower (diffusion-based video/action output), eliminating the need for multiple separate models. Available in two sizes — Cosmos 3 Nano (8B) and Cosmos 3 Super (32B) — with fully open model weights, training scripts, deployment tools, and six synthetic datasets on Hugging Face.

Unified Mixture-of-Transformers (MoT) architecture — Cosmos 3 combines a Reasoner tower (autoregressive VLM that interprets images, video, and text to understand motion and physical context) with a Generator tower (diffusion-based process for physics-aware video and action output); the reasoner can run independently, while the generator activates both towers — eliminating multi-model orchestration and simplifying physical AI development pipelines
Two model sizes for every deployment scenario — Cosmos 3 Nano (8B parameters) is optimized for workstation-grade inference on NVIDIA RTX PRO 6000 for real-time robotics; Cosmos 3 Super (32B parameters) targets datacenter Hopper/Blackwell GPUs for large-scale synthetic data generation; both available as NVIDIA NIM microservices for production-ready deployment without manual infra tuning
Open-source SOTA across physical AI benchmarks — leads on PAIBench-G, R-Bench, Physics-IQ, RoboLab, VANTAGE-Bench (reasoning), and Artificial Analysis Text-to-Image and Image-to-Video leaderboards; ships with fully open training recipes (SFT + action post-training), six synthetic datasets covering robotics/AV/warehouses/physics/human motion, and the new Cosmos HUE evaluation framework for rigorous video generation quality assessment
PricingOpen Source (Apache 2.0) + NIM API
RegionGlobal
2026NVIDIA

NemoClaw for LangChain Deep Agents Code v0.0.87

Open Source
Featured
Coding Agent
Software Engineering

NemoClaw for LangChain Deep Agents Code is a governed blueprint that allows developers to run open-source coding agents using NVIDIA Nemotron 3 Ultra models. It provides a secure, audit-friendly environment for performing complex software engineering tasks like refactoring, dependency upgrades, and automated code maintenance.

Governed OpenShell runtime environment with deny-by-default networking
Integration of NVIDIA Nemotron 3 Ultra with LangChain Deep Agents Code
Support for multi-step engineering tasks like refactoring and dependency upgrades
PricingFree & Open-source
Version0.0.87
2026NVIDIA

Nemotron-RL Agentic Terminal Pivot v1

Open Source
agentic
code

Nemotron-RL Agentic Terminal Pivot v1 is an open-source reinforcement learning dataset designed to train large language models for agentic command-line interface (CLI) tasks. It provides high-quality, expert-derived trajectories that enable models to perform complex software engineering operations, tool use, and reasoning within Linux environments.

Expert-derived trajectories for command-line agent training
Supports reinforcement learning from verifiable reward (RLVR)
Optimized for software engineering and tool-use reasoning
PricingFree
Version1
2025NVIDIA

Nemotron 3 Nano v3 Nano

Open Source
AI Agents
Agentic AI

Nemotron 3 Nano is an open-source language model designed for agentic AI applications, featuring a Mixture of Experts hybrid Mamba Transformer architecture with approximately 31.6 billion parameters. It offers efficient long-context reasoning capabilities, supporting up to 1 million tokens, and is optimized for multi-agent systems operating on extensive documents and codebases.

Mixture of Experts hybrid Mamba Transformer architecture
Approximately 31.6 billion parameters with 3.2 billion active per forward pass
Supports context lengths up to 1 million tokens
PricingFree
Version3 Nano
2026NVIDIA

NeMo Agent Toolkit v1.4

Open Source
Agentic AI
AI Agents

NVIDIA NeMo™ Agent Toolkit is an open-source AI framework designed to build, profile, and optimize AI agents and tools across various frameworks. It enables unified, cross-framework integration for connected AI agent systems, helping enterprises efficiently scale agentic systems while maintaining reliability. ([developer.nvidia.com](https://developer.nvidia.com/agent-intelligence-toolkit?utm_source=openai))

Framework agnostic, compatible with existing agentic frameworks like LangChain, LlamaIndex, CrewAI, and Microsoft Semantic Kernel.
Offers profiling and observability tools to monitor and debug workflows, integrating with platforms such as Phoenix, Weave, and Langfuse.
Provides an evaluation system to validate and maintain the accuracy of agentic workflows.
PricingFree
Version1.4
2026NVIDIA

NVIDIA Nemotron 3 Ultra v3 Ultra

Open Source
AI
Machine Learning

NVIDIA Nemotron 3 Ultra is a 550-billion-parameter Mixture-of-Experts model designed to enhance the efficiency and speed of long-running agents in complex workflows. It combines advanced reasoning capabilities with high throughput, enabling agents to perform tasks faster and more cost-effectively.

550-billion-parameter Mixture-of-Experts model
Hybrid Mamba-Transformer layers for efficient long-context handling
NVFP4 quantization for cross-architecture GPU deployment with up to 5x higher throughput
PricingFree
Version3 Ultra

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode