Introducing Unsloth Studio: a new web UI for local AI
NVIDIA releases Nemotron-3-Nano-4B, a 4B open hybrid MoE model that follows Nemotron-3-Super-120B-A12B and Nemotron-3-Nano-30B-A3B. The Nemotron family is designed for fast, accurate coding, math, and agentic workloads. They feature a 1M-token context window and are competitive across reasoning, chat, and throughput benchmarks.
Nemotron-3-Nano-4B runs on 5GB of RAM, VRAM, or unified memory. Nemotron-3-Nano-30A3B runs on 24GB RAM. Nemotron 3 can now be fine-tuned locally via Unsloth. Thanks to NVIDIA for giving Unsloth day-zero support.
Usage Guide
NVIDIA recommends these settings for inference:
General chat/instruction (default):
temperature = 1.0top_p = 1.0
Tool calling use-cases:
temperature = 0.6top_p = 0.95
For most local use, set:
max_new_tokens=32,768to262,144for standard prompts with a max of 1M tokens- Increase for deep reasoning or long-form generation as your RAM/VRAM allows.
The chat template format is found when we use the below:
tokenizer.apply_chat_template([
{"role" : "user", "content" : "What is 1+1?"},
{"role" : "assistant", "content" : "2"},
{"role" : "user", "content" : "What is 2+2?"}
], add_generation_prompt = True, tokenize = False,
)
Because the model was trained with NoPE, you only need to change max_position_embeddings. The model doesn’t use explicit positional embeddings, so YaRN isn’t needed.
Nemotron 3 chat template format:
Nemotron 3 uses <think> with token ID 12 and </think> with token ID 13 for reasoning. Use --special to see the tokens for llama.cpp. You might also need --verbose-prompt to see <think> since it's prepended.
<|im_start|>system\n<|im_end|>\n<|im_start|>user\nWhat is 1+1?<|im_end|>\n<|im_start|>assistant\n<think></think>2<|im_end|>\n<|im_start|>user\nWhat is 2+2?<|im_end|>\n<|im_start|>assistant\n<think>\n```
## Run Nemotron-3-Nano-4B
Depending on your use-case you will need to use different settings.
### Unsloth Studio Guide
Nemotron 3 can be run and fine-tuned in [Unsloth Studio](/content/docs/new/studio/index.html), our new open-source web UI for local AI. With Unsloth Studio, you can run models locally on **MacOS, Windows**, Linux and:
- Search, download, [run GGUFs](/content/docs/new/studio#run-models-locally/index.html) and safetensor models
- [**Self-healing** tool calling](/content/docs/new/studio#execute-code--heal-tool-calling/index.html)
- [**Code execution**](/content/docs/new/studio#run-models-locally/index.html)(Python, Bash)
- Automatic inference parameter tuning (temp, top-p, etc.)
- Fast CPU + GPU inference via llama.cpp
- [Train LLMs](/content/docs/new/studio#no-code-training/index.html) 2x faster with 70% less VRAM
### Install Unsloth
Run in your terminal:
#### MacOS, Linux, WSL:
```bash
curl -fsSL https://unsloth.ai/install.sh | sh
Windows PowerShell:
irm https://unsloth.ai/install.ps1 | iex
Launch Unsloth
MacOS, Linux, WSL, Windows:
unsloth studio -H 0.0.0.0 -p 8888
Then open http://localhost:8888 in your browser.
Run Nemotron-3-Nano-30B-A3B
Depending on your use-case you will need to use different settings.
Llama.cpp Tutorial:
Instructions to run in llama.cpp (we'll be using 8-bit for near full precision):
- Obtain the latest
llama.cppon GitHub here. - You can directly pull from Hugging Face. You can increase the context to 1M as your RAM/VRAM allows.
- Download the model via:
import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from huggingface_hub import snapshot_download
snapshot_download(
repo_id = "unsloth/Nemotron-3-Nano-30B-A3B-GGUF",
local_dir = "unsloth/Nemotron-3-Nano-30B-A3B-GGUF",
allow_patterns = ["*UD-Q4_K_XL*"],
)
- Also, adjust context window as required. Ensure your hardware can handle more than a 256K context window. Setting it to 1M may trigger CUDA OOM and crash, which is why the default is 262,144.
Fine-tuning Nemotron 3 and RL
Unsloth now supports fine-tuning of all Nemotron models, including Nemotron 3 Super and Nano. The 4B model fits on a free Colab GPU however the 30B model does not fit. We still made an 80GB A100 Colab notebook for you to fine-tune with. 16-bit LoRA fine-tuning of Nemotron 3 Nano will use around 60GB VRAM:
Reinforcement Learning + NeMo Gym
We worked with the open-source NVIDIA NeMo Gym team to enable the democratization of RL environments.
Llama-server serving & deployment
To deploy Nemotron 3 for production, we use llama-server In a new terminal say via tmux, deploy the model via:
./llama.cpp/llama-server \
--model unsloth/Nemotron-3-Nano-30B-A3B-GGUF/Nemotron-3-Nano-30B-A3B-UD-Q4_K_XL.gguf \
--alias "unsloth/Nemotron-3-Nano-30B-A3B" \
--prio 3 \
--min_p 0.01 \
--temp 0.6 \
--top-p 0.95 \
--ctx-size 16384 \
--port 8001
Benchmarks
Nemotron-3-Nano-4B is the best performing model for its size, including throughput.
Nemotron-3-Nano-30B-A3B is the best performing model across all benchmarks, including throughput.