Introducing Unsloth Studio: a new web UI for local AI

NVIDIA releases Nemotron-3-Nano-4B, a 4B open hybrid MoE model that follows Nemotron-3-Super-120B-A12B and Nemotron-3-Nano-30B-A3B. The Nemotron family is designed for fast, accurate coding, math, and agentic workloads. They feature a 1M-token context window and are competitive across reasoning, chat, and throughput benchmarks.

Nemotron-3-Nano-4B runs on 5GB of RAM, VRAM, or unified memory. Nemotron-3-Nano-30A3B runs on 24GB RAM. Nemotron 3 can now be fine-tuned locally via Unsloth. Thanks to NVIDIA for giving Unsloth day-zero support.

Usage Guide

NVIDIA recommends these settings for inference:

General chat/instruction (default):

  • temperature = 1.0
  • top_p = 1.0

Tool calling use-cases:

  • temperature = 0.6
  • top_p = 0.95

For most local use, set:

  • max_new_tokens = 32,768 to 262,144 for standard prompts with a max of 1M tokens
  • Increase for deep reasoning or long-form generation as your RAM/VRAM allows.

The chat template format is found when we use the below:


tokenizer.apply_chat_template([
    {"role" : "user", "content" : "What is 1+1?"},
    {"role" : "assistant", "content" : "2"},
    {"role" : "user", "content" : "What is 2+2?"}
    ], add_generation_prompt = True, tokenize = False,
)

Because the model was trained with NoPE, you only need to change max_position_embeddings. The model doesn’t use explicit positional embeddings, so YaRN isn’t needed.

Nemotron 3 chat template format:

Nemotron 3 uses <think> with token ID 12 and </think> with token ID 13 for reasoning. Use --special to see the tokens for llama.cpp. You might also need --verbose-prompt to see <think> since it's prepended.

<|im_start|>system\n<|im_end|>\n<|im_start|>user\nWhat is 1+1?<|im_end|>\n<|im_start|>assistant\n<think></think>2<|im_end|>\n<|im_start|>user\nWhat is 2+2?<|im_end|>\n<|im_start|>assistant\n<think>\n```

## Run Nemotron-3-Nano-4B

Depending on your use-case you will need to use different settings.

### Unsloth Studio Guide

Nemotron 3 can be run and fine-tuned in [Unsloth Studio](/content/docs/new/studio/index.html), our new open-source web UI for local AI. With Unsloth Studio, you can run models locally on **MacOS, Windows**, Linux and:
- Search, download, [run GGUFs](/content/docs/new/studio#run-models-locally/index.html) and safetensor models
- [**Self-healing** tool calling](/content/docs/new/studio#execute-code--heal-tool-calling/index.html) 
- [**Code execution**](/content/docs/new/studio#run-models-locally/index.html)(Python, Bash)
- Automatic inference parameter tuning (temp, top-p, etc.)
- Fast CPU + GPU inference via llama.cpp
- [Train LLMs](/content/docs/new/studio#no-code-training/index.html) 2x faster with 70% less VRAM

### Install Unsloth

Run in your terminal:
#### MacOS, Linux, WSL:
```bash
curl -fsSL https://unsloth.ai/install.sh | sh

Windows PowerShell:

irm https://unsloth.ai/install.ps1 | iex

Launch Unsloth

MacOS, Linux, WSL, Windows:

unsloth studio -H 0.0.0.0 -p 8888

Then open http://localhost:8888 in your browser.

Run Nemotron-3-Nano-30B-A3B

Depending on your use-case you will need to use different settings.

Llama.cpp Tutorial:

Instructions to run in llama.cpp (we'll be using 8-bit for near full precision):

  1. Obtain the latest llama.cpp on GitHub here.
  2. You can directly pull from Hugging Face. You can increase the context to 1M as your RAM/VRAM allows.
  3. Download the model via:
import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from huggingface_hub import snapshot_download
snapshot_download(
    repo_id = "unsloth/Nemotron-3-Nano-30B-A3B-GGUF",
    local_dir = "unsloth/Nemotron-3-Nano-30B-A3B-GGUF",
    allow_patterns = ["*UD-Q4_K_XL*"],
)
  1. Also, adjust context window as required. Ensure your hardware can handle more than a 256K context window. Setting it to 1M may trigger CUDA OOM and crash, which is why the default is 262,144.

Fine-tuning Nemotron 3 and RL

Unsloth now supports fine-tuning of all Nemotron models, including Nemotron 3 Super and Nano. The 4B model fits on a free Colab GPU however the 30B model does not fit. We still made an 80GB A100 Colab notebook for you to fine-tune with. 16-bit LoRA fine-tuning of Nemotron 3 Nano will use around 60GB VRAM:

Reinforcement Learning + NeMo Gym

We worked with the open-source NVIDIA NeMo Gym team to enable the democratization of RL environments.

Llama-server serving & deployment

To deploy Nemotron 3 for production, we use llama-server In a new terminal say via tmux, deploy the model via:

./llama.cpp/llama-server \
    --model unsloth/Nemotron-3-Nano-30B-A3B-GGUF/Nemotron-3-Nano-30B-A3B-UD-Q4_K_XL.gguf \
    --alias "unsloth/Nemotron-3-Nano-30B-A3B" \
    --prio 3 \
    --min_p 0.01 \
    --temp 0.6 \
    --top-p 0.95 \
    --ctx-size 16384 \
    --port 8001

Benchmarks

Nemotron-3-Nano-4B is the best performing model for its size, including throughput.

Nemotron-3-Nano-30B-A3B is the best performing model across all benchmarks, including throughput.