# Introducing Unsloth Studio: a new web UI for local AI

NVIDIA releases **Nemotron-3-Nano-4B**, a 4B open hybrid MoE model that follows [Nemotron-3-Super-120B-A12B](/content/docs/models/nemotron-3/nemotron-3-super/index.html) and Nemotron-3-Nano-30B-A3B. The Nemotron family is designed for fast, accurate coding, math, and agentic workloads. They feature a **1M-token context** window and are competitive across reasoning, chat, and throughput benchmarks.

Nemotron-3-Nano-4B runs on **5GB** of RAM, VRAM, or unified memory. Nemotron-3-Nano-30A3B runs on **24GB** RAM. Nemotron 3 can now be fine-tuned locally via [Unsloth](https://github.com/unslothai/unsloth). Thanks to NVIDIA for giving Unsloth day-zero support.

### Usage Guide

NVIDIA recommends these settings for inference:

**General chat/instruction (default):**
- `temperature = 1.0`
- `top_p = 1.0`

**Tool calling use-cases:**
- `temperature = 0.6`
- `top_p = 0.95`

**For most local use, set:**
- `max_new_tokens` = `32,768` to `262,144` for standard prompts with a max of 1M tokens
- Increase for deep reasoning or long-form generation as your RAM/VRAM allows.

The chat template format is found when we use the below:

```python

tokenizer.apply_chat_template([
    {"role" : "user", "content" : "What is 1+1?"},
    {"role" : "assistant", "content" : "2"},
    {"role" : "user", "content" : "What is 2+2?"}
    ], add_generation_prompt = True, tokenize = False,
)
```

Because the model was trained with NoPE, you only need to change `max_position_embeddings`. The model doesn’t use explicit positional embeddings, so YaRN isn’t needed.

#### Nemotron 3 chat template format:

Nemotron 3 uses `<think>` with token ID 12 and `</think>` with token ID 13 for reasoning. Use `--special` to see the tokens for llama.cpp. You might also need `--verbose-prompt` to see `<think>` since it's prepended.

```python
<|im_start|>system\n<|im_end|>\n<|im_start|>user\nWhat is 1+1?<|im_end|>\n<|im_start|>assistant\n<think></think>2<|im_end|>\n<|im_start|>user\nWhat is 2+2?<|im_end|>\n<|im_start|>assistant\n<think>\n```

## Run Nemotron-3-Nano-4B

Depending on your use-case you will need to use different settings.

### Unsloth Studio Guide

Nemotron 3 can be run and fine-tuned in [Unsloth Studio](/content/docs/new/studio/index.html), our new open-source web UI for local AI. With Unsloth Studio, you can run models locally on **MacOS, Windows**, Linux and:
- Search, download, [run GGUFs](/content/docs/new/studio#run-models-locally/index.html) and safetensor models
- [**Self-healing** tool calling](/content/docs/new/studio#execute-code--heal-tool-calling/index.html) 
- [**Code execution**](/content/docs/new/studio#run-models-locally/index.html)(Python, Bash)
- Automatic inference parameter tuning (temp, top-p, etc.)
- Fast CPU + GPU inference via llama.cpp
- [Train LLMs](/content/docs/new/studio#no-code-training/index.html) 2x faster with 70% less VRAM

### Install Unsloth

Run in your terminal:
#### MacOS, Linux, WSL:
```bash
curl -fsSL https://unsloth.ai/install.sh | sh
```
#### Windows PowerShell:
```bash
irm https://unsloth.ai/install.ps1 | iex
```

### Launch Unsloth
**MacOS, Linux, WSL, Windows:**
```bash
unsloth studio -H 0.0.0.0 -p 8888
```
**Then open** `http://localhost:8888` **in your browser.**

### Run Nemotron-3-Nano-30B-A3B

Depending on your use-case you will need to use different settings.

### Llama.cpp Tutorial:
Instructions to run in llama.cpp (we'll be using 8-bit for near full precision):
1. Obtain the latest `llama.cpp` on [GitHub here](https://github.com/ggml-org/llama.cpp).
2. You can directly pull from Hugging Face. You can increase the context to 1M as your RAM/VRAM allows.
3. Download the model via:
```python
import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from huggingface_hub import snapshot_download
snapshot_download(
    repo_id = "unsloth/Nemotron-3-Nano-30B-A3B-GGUF",
    local_dir = "unsloth/Nemotron-3-Nano-30B-A3B-GGUF",
    allow_patterns = ["*UD-Q4_K_XL*"],
)
```
4. Also, adjust **context window** as required. Ensure your hardware can handle more than a 256K context window. Setting it to 1M may trigger CUDA OOM and crash, which is why the default is 262,144.

### Fine-tuning Nemotron 3 and RL

Unsloth now supports fine-tuning of all Nemotron models, including Nemotron 3 Super and Nano. The 4B model fits on a free Colab GPU however the 30B model does not fit. We still made an 80GB A100 Colab notebook for you to fine-tune with. 16-bit LoRA fine-tuning of Nemotron 3 Nano will use around **60GB VRAM**:
- [Nemotron-3-Nano-30B-A3B SFT LoRA notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Nemotron-3-Nano-30B-A3B_A100.ipynb)

### Reinforcement Learning + NeMo Gym

We worked with the open-source NVIDIA [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym/pull/492) team to enable the democratization of RL environments.

### Llama-server serving & deployment
To deploy Nemotron 3 for production, we use `llama-server` In a new terminal say via tmux, deploy the model via:
```bash
./llama.cpp/llama-server \
    --model unsloth/Nemotron-3-Nano-30B-A3B-GGUF/Nemotron-3-Nano-30B-A3B-UD-Q4_K_XL.gguf \
    --alias "unsloth/Nemotron-3-Nano-30B-A3B" \
    --prio 3 \
    --min_p 0.01 \
    --temp 0.6 \
    --top-p 0.95 \
    --ctx-size 16384 \
    --port 8001
```

### Benchmarks
Nemotron-3-Nano-4B is the best performing model for its size, including throughput.

Nemotron-3-Nano-30B-A3B is the best performing model across all benchmarks, including throughput.
