Nemotron 3 Super Tutorial: Multi-Token Prediction, Latent MoE, Perplexity and OpenCode Integration

NVIDIA Developer · Intermediate ·🧠 Large Language Models ·4mo ago

Key Takeaways

The video demonstrates the NVIDIA Neotron 3 Super model, a hybrid Mamba-Transformer MoE with multi-token prediction and latent MoE, and shows how to use it with OpenCode and Perplexity integration.

Full Transcript

Hey, what's up everybody? My name is Chris and I'm really excited to talk to you today about our brand new Neotron 3 super model, the latest entrant in our Neotron 3 family of models. Once again, this is a hybrid Mamba transformer architecture. Uh we've added a few bells and whistles though, so it has multi-token prediction as well as latente. You can read more about those things in the technical report, but today we're going to get hands-on and see how we can actually use the model. It's a 120B total, 12B active parameter model. Uh, which means that it can still rip pretty fast. So, what we're going to do is look at how we can access it through build.vidia.com, look a little bit about the integration that we have with Perplexity, uh, look through some code in a notebook, and then showcase how you can use the model with something like Open Code as your driver for your agent harness. All right, let's dive right in, and we'll see what this model can do. All right, the first thing we can do is check out the model on uh build.unvideo.com. So you can see here we have our Neotron 3 Super 12B- A12B. Uh we can ask it questions like, "Hey Super, how's it going?" We're going to get a reasoning trace, of course, and then the final answer. We can do things like ask it a riddle. And this is basically the standard interface for using the model, right? Uh what we can get from here though is actually our API key by clicking this generate API key button. So what we're going to do is generate that key and then head to the notebook. All right. So now that we have our API key, we can work through this simple walkthrough. Uh we're going to look at how we can enable thinking, how we can leverage reasoning budget, and how we can use the brand new loweffort reasoning mode that's unique to Neotron 3 Super. We've already got our API key, so we can skip these instructions just fine. We're going to set up our base URL as integrate.api.invidia.com/b1 just like the instructions said. And we're going to have our model be exactly what we just looked at at build.envidia.com. First things first, let's look at the model in its default mode, which is reasoning on. You'll notice that I explicitly passed in enable thinking true here. But even if I didn't pass this parameter, our model would actually operate the same as we saw before. So let's go ahead and try that like this. And you can see that we get the uh default behavior which is a little bit of thinking and then the response. Now we can use a reasoning budget in order to tell the model how long to think for. What this is going to do is using enable thinking true allows us to you know set a maximum token budget. So if our model needed to think for more than 8,000 tokens, we would truncate its thinking and start returning the answer. This is great if you want a really granular control of the model's reasoning. Next, we have loweffort mode. Loweffort mode is what it sounds like. The model just tries to think as little as possible. Uh this is something that we can set just by setting loweffort equal to true. Uh the recommendation would be to start with the normal behavior of the model. If it's a little bit too ver verbose, you can move to using uh the loweffort flag and then if you need something in between at that point experiment with different potential reasoning budgets. We also have an integration with perplexity which actually allows us to use this model along with their web search tool through perplexity. So we can ask a question like what is the new architecture for neotron 3 family and get this awesome response as found by the search tool through perplexity. Not only can we use it in the notebook like this, but we also have access to it through the model dropdown in perplexity.ai. So we can actually select super as our model and ask it to look up something. So we can use it to look up something like what is the newest architecture for the Neotron 3 family of models and then send this request off. And same thing as before, it's going to think, look up, review, and then provide us with a final response. Very cool. You'll notice in the notebook that we also have this open code blob. This open code blob is actually going to let us use the model with open code. So let's look at how we would do that. All right. So we can head to our terminal. And what we're going to do is go ahead and open up our open.json. I have a demo version just to hide my API key. But the idea is that we can set our schema model and our provider. For our provider, we're going to call it NVIDIA NIM. We're going to include our base URL, which is the same URL that we saw in the notebook, as well as our model, which is the same model we saw at build.vidia.com. Once we have this set up, we're good to go ahead and actually run open code, which we can do with the command open code. And then we're going to find our model in the models section. We can search it up with super and see that we have our old versions of super as well as the latest Neatron 3 Super. Now that we've picked our model, we can go ahead and ask it to do stuff for us, like perhaps, you know, build a cool snake game in HTML CSS. It's going to start cooking. And that's it. And finally, we look at some of the things that our open code agent was able to build, like this sweet landing page with animations, cool things when we hover a slide. Uh, so this is for a slide deck. And of course, you know, we got to add the snake game. So, we have this snake game that we can play, which keeps track of our high score, as well as, of course, shows us a fun little animation whenever we hit a fruit. Thanks so much for watching this tutorial video. I hope you enjoyed it. We have some resources prepared that you can check out down below, including a tech blog, a tech report, as well as a number of examples on how you can actually customize this model if you wanted to do that. Thanks so much for watching and have a great

Original Description

NVIDIA Nemotron 3 Super is a 120B total, 12B active-parameter hybrid Mamba-Transformer MoE designed for high-efficiency agentic reasoning. It’s a useful model for agentic and coding tasks - with its 1M context window making it a great option to power agent harnesses like OpenCode. 🛠️ Key Technical Highlights: - Hybrid Backbone: Interleaving Mamba-2 for sequence efficiency and Transformer layers for precision reasoning. - Latent MoE: Routing compressed tokens to 4x as many experts for the same inference cost. - Native NVFP4: Pretrained specifically for NVIDIA Blackwell to cut memory requirements and speed up inference. - API Control: Implementation of enable_thinking, reasoning_budget, and low_effort modes for granular control. NVIDIA Technical Blog: https://nvda.ws/47oQ6jX NVIDIA Tech Report: https://nvda.ws/4cxO98m Nemotron 3 Super on Hugging Face: https://nvda.ws/3ORn5at 0:00 - Intro to Nemotron-3 Super (120B/12B) 0:55 - Live Demo: Reasoning & Riddles 1:28 - Python Tutorial: API Key & Setup 1:56 - Customizing "Thinking" & Reasoning Budgets 2:44 - Low-Effort vs. Deep Reasoning 3:12 - Nemotron-3 Super on Perplexity AI 3:57 - Building with OpenCode (HTML/CSS Landing Pages) 5:15 - The Neon Snake Game Reveal 5:30 - Resources & Wrap-up
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from NVIDIA Developer · NVIDIA Developer · 0 of 60

← Previous Next →
1 Ray Tracing Essentials Part 2: Rasterization versus Ray Tracing
Ray Tracing Essentials Part 2: Rasterization versus Ray Tracing
NVIDIA Developer
2 Ray Tracing Essentials Part 3: Ray Tracing Hardware
Ray Tracing Essentials Part 3: Ray Tracing Hardware
NVIDIA Developer
3 Ray Tracing Essentials Part 4: The Ray Tracing Pipeline
Ray Tracing Essentials Part 4: The Ray Tracing Pipeline
NVIDIA Developer
4 NsightGraphics 2020 2 Release Spotlight
NsightGraphics 2020 2 Release Spotlight
NVIDIA Developer
5 Ray Tracing Essentials Part 5: Ray Tracing Effects
Ray Tracing Essentials Part 5: Ray Tracing Effects
NVIDIA Developer
6 Ray Tracing Essentials Part 6: The Rendering Equation
Ray Tracing Essentials Part 6: The Rendering Equation
NVIDIA Developer
7 Ray Tracing Essentials Part 7: Denoising for Ray Tracing
Ray Tracing Essentials Part 7: Denoising for Ray Tracing
NVIDIA Developer
8 Spatiotemporal Importance Resampling for Many-Light Ray Tracing (ReSTIR)
Spatiotemporal Importance Resampling for Many-Light Ray Tracing (ReSTIR)
NVIDIA Developer
9 Announcing Cloud-Native Support for Jetson Platform
Announcing Cloud-Native Support for Jetson Platform
NVIDIA Developer
10 JetsonTV: Build your next project with NVIDIA Jetson
JetsonTV: Build your next project with NVIDIA Jetson
NVIDIA Developer
11 Nsight Compute Feature Spotlight: Roofline Analysis, Asynchronous Copy, Sparse Data Compression
Nsight Compute Feature Spotlight: Roofline Analysis, Asynchronous Copy, Sparse Data Compression
NVIDIA Developer
12 Nsight Systems Feature Spotlight: OpenMP
Nsight Systems Feature Spotlight: OpenMP
NVIDIA Developer
13 Isaac Sim 2020: Deep Dive
Isaac Sim 2020: Deep Dive
NVIDIA Developer
14 NVIDIA Jetson: Enabling AI-Powered Autonomous Machines at Scale
NVIDIA Jetson: Enabling AI-Powered Autonomous Machines at Scale
NVIDIA Developer
15 NVIDIA Tools to Train, Build, and Deploy Intelligent Vision Applications at the Edge
NVIDIA Tools to Train, Build, and Deploy Intelligent Vision Applications at the Edge
NVIDIA Developer
16 Jetson Xavier NX Developer Kit: The Next Leap in Edge Computing
Jetson Xavier NX Developer Kit: The Next Leap in Edge Computing
NVIDIA Developer
17 Synthesizing High-Resolution Images with StyleGAN2
Synthesizing High-Resolution Images with StyleGAN2
NVIDIA Developer
18 NVIDIA Robotics: Isaac SDK and Sim 2020.1
NVIDIA Robotics: Isaac SDK and Sim 2020.1
NVIDIA Developer
19 Accelerating COVID-19 Research with GPUs
Accelerating COVID-19 Research with GPUs
NVIDIA Developer
20 Visualizing 150 Terabytes of Data
Visualizing 150 Terabytes of Data
NVIDIA Developer
21 Boosting Performance and Utilization with Multi-Instance GPU
Boosting Performance and Utilization with Multi-Instance GPU
NVIDIA Developer
22 Running Multiple Workloads on a Single A100 GPU
Running Multiple Workloads on a Single A100 GPU
NVIDIA Developer
23 NVIDIA Nsight Feature Spotlight: GPU Trace
NVIDIA Nsight Feature Spotlight: GPU Trace
NVIDIA Developer
24 Spark 3 Demo: Comparing Performance of GPUs vs. CPUs
Spark 3 Demo: Comparing Performance of GPUs vs. CPUs
NVIDIA Developer
25 NVIDIA Jetson Nano Wins Edge AI and Vision Alliance Award
NVIDIA Jetson Nano Wins Edge AI and Vision Alliance Award
NVIDIA Developer
26 NVIDIA IndeX on Google Cloud Platform Marketplace
NVIDIA IndeX on Google Cloud Platform Marketplace
NVIDIA Developer
27 DeepStream SDK: Best practices for performance optimization
DeepStream SDK: Best practices for performance optimization
NVIDIA Developer
28 Efficiently Deploying GPU Accelerated 5G CloudRAN for Edge AI Inferencing
Efficiently Deploying GPU Accelerated 5G CloudRAN for Edge AI Inferencing
NVIDIA Developer
29 NVIDIA PhysicsNeMo - Accelerating Scientific & Engineering Simulation Workflows with AI
NVIDIA PhysicsNeMo - Accelerating Scientific & Engineering Simulation Workflows with AI
NVIDIA Developer
30 NVIDIA Deep Learning Institute Instructor-Led Training Available Remotely
NVIDIA Deep Learning Institute Instructor-Led Training Available Remotely
NVIDIA Developer
31 Advancing AR Glasses
Advancing AR Glasses
NVIDIA Developer
32 Blender Cycles: RTX On
Blender Cycles: RTX On
NVIDIA Developer
33 Real-Time GPU-Accelerated Data Analytics of 250 million Flight Data Records of 737 Max grounding
Real-Time GPU-Accelerated Data Analytics of 250 million Flight Data Records of 737 Max grounding
NVIDIA Developer
34 Assessing Property Damage with AI
Assessing Property Damage with AI
NVIDIA Developer
35 RAPIDS: GPU-Accelerated Data Analytics & Machine Learning
RAPIDS: GPU-Accelerated Data Analytics & Machine Learning
NVIDIA Developer
36 DaVinci Resolve Turns RTX On
DaVinci Resolve Turns RTX On
NVIDIA Developer
37 RAPIDS with Plotly Dash : GPU-Accelerated Census 2010 Visualization
RAPIDS with Plotly Dash : GPU-Accelerated Census 2010 Visualization
NVIDIA Developer
38 NVIDIA IndeX for arivis5D Cloud Platform
NVIDIA IndeX for arivis5D Cloud Platform
NVIDIA Developer
39 NVIDIA Backchannel: Behind the Scenes of Marbles at Night RTX
NVIDIA Backchannel: Behind the Scenes of Marbles at Night RTX
NVIDIA Developer
40 NVIDIA Backchannel: Sneak Peek into Marbles RTX in Omniverse
NVIDIA Backchannel: Sneak Peek into Marbles RTX in Omniverse
NVIDIA Developer
41 How to Create "Paint" in Substance Painter
How to Create "Paint" in Substance Painter
NVIDIA Developer
42 Accelerate AI development for Computer Vision on the NVIDIA Jetson with alwaysAI
Accelerate AI development for Computer Vision on the NVIDIA Jetson with alwaysAI
NVIDIA Developer
43 Securing Next Generation Apps over VMware Cloud Foundation with Bluefield-2 DPU
Securing Next Generation Apps over VMware Cloud Foundation with Bluefield-2 DPU
NVIDIA Developer
44 Accelerated Data Centers with NVIDIA and VMware
Accelerated Data Centers with NVIDIA and VMware
NVIDIA Developer
45 GPU-Accelerated Motion Blur in Blender Cycles
GPU-Accelerated Motion Blur in Blender Cycles
NVIDIA Developer
46 NVIDIA Clara Guardian Virtual Patient Assistant
NVIDIA Clara Guardian Virtual Patient Assistant
NVIDIA Developer
47 Revolutionizing Supercomputing with NVIDIA UFM Cyber-AI
Revolutionizing Supercomputing with NVIDIA UFM Cyber-AI
NVIDIA Developer
48 Inventing Virtual Meetings of Tomorrow with NVIDIA AI Research
Inventing Virtual Meetings of Tomorrow with NVIDIA AI Research
NVIDIA Developer
49 Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
NVIDIA Developer
50 Getting started with Jetson Nano 2GB Developer Kit
Getting started with Jetson Nano 2GB Developer Kit
NVIDIA Developer
51 NVIDIA Jetson Developer Community AI Projects
NVIDIA Jetson Developer Community AI Projects
NVIDIA Developer
52 Open-source projects on NVIDIA Jetson Nano 2GB Developer Kit
Open-source projects on NVIDIA Jetson Nano 2GB Developer Kit
NVIDIA Developer
53 Real-Time Ray Tracing with Project Lavina
Real-Time Ray Tracing with Project Lavina
NVIDIA Developer
54 Jetson AI Fundamentals - S1E2 - Hello Camera
Jetson AI Fundamentals - S1E2 - Hello Camera
NVIDIA Developer
55 Develop Optimized Conversational AI Models with NVIDIA NeMo on DGX A100
Develop Optimized Conversational AI Models with NVIDIA NeMo on DGX A100
NVIDIA Developer
56 Jetson AI Fundamentals - S1E4 - Image Regression Project
Jetson AI Fundamentals - S1E4 - Image Regression Project
NVIDIA Developer
57 Jetson AI Fundamentals - S2E1 - JetBot Intro and Hardware
Jetson AI Fundamentals - S2E1 - JetBot Intro and Hardware
NVIDIA Developer
58 Jetson AI Fundamentals - S2E2 - JetBot Software Setup
Jetson AI Fundamentals - S2E2 - JetBot Software Setup
NVIDIA Developer
59 Jetson AI Fundamentals - S1E1 - First Time Setup with JetPack
Jetson AI Fundamentals - S1E1 - First Time Setup with JetPack
NVIDIA Developer
60 Jetson AI Fundamentals - S1E3 - Image Classification Project
Jetson AI Fundamentals - S1E3 - Image Classification Project
NVIDIA Developer

This video tutorial demonstrates the NVIDIA Neotron 3 Super model and shows how to use it with OpenCode and Perplexity integration. It covers topics such as multi-token prediction, latent MoE, and reasoning budget.

Key Takeaways
  1. Access the Neotron 3 Super model on build.vidia.com
  2. Generate an API key
  3. Use the API key in a notebook to leverage the model's capabilities
  4. Experiment with reasoning budget and loweffort mode
  5. Integrate the model with OpenCode and Perplexity
💡 The Neotron 3 Super model can be used for agentic reasoning and coding tasks, and its integration with OpenCode and Perplexity enables more efficient and effective use cases.

Related Reads

📰
Introducing Claude Opus 5 on AWS: Anthropic’s most capable Opus model
Learn about Claude Opus 5, Anthropic's most capable Opus model, and how to integrate it into agentic systems on AWS
AWS Machine Learning
📰
Can AI Keep a Great Mind Alive?
Learn how AI can preserve a great mind's thinking through persona fine-tuning, first-principles reasoning, and mechanistic interpretability
Dev.to AI
📰
Anthropic launches Opus 5
Anthropic's Opus 5 offers a cheaper and less restrictive alternative to Fable, making it a preferable choice for most use cases
TechCrunch AI
📰
Claude Opus 5 arrives with near Fable performance at half the price
Learn about Claude Opus 5's upgraded features and improved performance at a lower price point, making it an attractive option for developers and enterprises
ZDNet

Chapters (9)

Intro to Nemotron-3 Super (120B/12B)
0:55 Live Demo: Reasoning & Riddles
1:28 Python Tutorial: API Key & Setup
1:56 Customizing "Thinking" & Reasoning Budgets
2:44 Low-Effort vs. Deep Reasoning
3:12 Nemotron-3 Super on Perplexity AI
3:57 Building with OpenCode (HTML/CSS Landing Pages)
5:15 The Neon Snake Game Reveal
5:30 Resources & Wrap-up
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →