# Introducing Unsloth Studio: a new web UI for local AI

Reinforcement learning's (RL) biggest challenge is supporting long reasoning traces. We're introducing new batching algorithms to enable ~ **7x longer context** (can be more than 12x) RL training with no accuracy or speed degradation vs. other optimized setups that use FA3, kernels & chunked losses.

- Unsloth now trains gpt-oss QLoRA with **380K context** on a single 192GB NVIDIA B200 GPU
- [Qwen3](/content/docs/models/tutorials/qwen3-how-to-run-and-fine-tune#fine-tuning-qwen3-with-unsloth/index.html)-8B GRPO reaches **110K context** on an 80GB VRAM H100 via [vLLM](/content/docs/get-started/reinforcement-learning-rl-guide/grpo-long-context#vllm-for-rl/index.html) and QLoRA, and **65K** for [gpt-oss](/content/docs/models/gpt-oss-how-to-run-and-fine-tune/gpt-oss-reinforcement-learning/index.html) with BF16 LoRA.
- On 24GB VRAM, gpt-oss reaches 20K context and 32K for [Qwen3-VL](/content/docs/models/tutorials/qwen3-how-to-run-and-fine-tune/qwen3-vl-how-to-run-and-fine-tune/index.html)-8B QLoRA
- Unsloth GRPO RL runs with Llama, Gemma & all models auto support longer contexts

Our new data-movement and batching kernels and algorithms unlock more context by:

- Dynamic [flattened sequence chunking](/content/docs/get-started/reinforcement-learning-rl-guide/grpo-long-context#flattened-sequence-length-chunking/index.html) to avoid materializing massive logit tensors
- [Offloading log softmax](/content/docs/get-started/reinforcement-learning-rl-guide/grpo-long-context#offloading-activations-for-log-softmax/index.html) activations which prevents silent memory growth over time.

**You can combine all features in Unsloth together:**

1. Unsloth's [weight-sharing](/content/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl/index.html) feature with [vLLM](https://github.com/vllm-project/vllm) and our Standby Feature in [Memory Efficient RL](/content/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl/index.html)
2. Unsloth's [Flex Attention](/content/docs/models/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training/index.html) for long context gpt-oss and our [500K Context Training](/content/docs/blog/500k-context-length-fine-tuning/index.html)
3. Float8 training in [FP8 RL](/content/docs/get-started/reinforcement-learning-rl-guide/fp8-reinforcement-learning/index.html) and Unsloth's [async gradient checkpointing](/content/blog/long-context/index.html) and much more

### 🎉Getting started

To get started, you can use any existing [GRPO notebooks](/content/docs/get-started/unsloth-notebooks#grpo-reasoning-rl-notebooks/index.html) (or update Unsloth if local):

[**gpt-oss-20b**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/gpt-oss-(20B)-GRPO.ipynb) 
[**Qwen3-VL-8B**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_VL_(8B)-Vision-GRPO.ipynb) 
[Qwen3-8B - **FP8**](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_8B_FP8_GRPO.ipynb)

Adopting Unsloth for your RL tasks provides a robust framework for managing large-scale models efficiently. To effectively utilize Unsloth's enhancements:

- **Hardware Recommendations**: Use of NVIDIA H100 or equivalent for optimal VRAM utilization.
- **Configuration Tips**: Ensure `batch_size` and `gradient_accumulation_steps` settings align with your computational resources for best performance.

Update Unsloth to the latest Pypi release to get the latest updates:

```bash
pip install --upgrade --no-cache-dir unsloth unsloth_zoo
```

### 🔢Flattened sequence length chunking

Previously, Unsloth reduced memory usage of RL by avoiding the full materialization of the logits tensor through chunking over the batch dimension. A rough estimate of the VRAM required to materialize logits during the forward pass is shown in Equation (1).

Equation 1: Logit Memory (GB)=batch size×context length×vocab dim/1024^3

Using this formulation, a configuration with `batch_size = 4`, `context_length = 8192`, and `vocab_dim = 128,000` would require approximately **3.3 GB of VRAM** to store the logits tensor.

Via [Long Context gpt-oss](/content/docs/models/gpt-oss-how-to-run-and-fine-tune/long-context-gpt-oss-training/index.html) last year, we then introduced a fused loss approach for GRPO. This approach ensures that only a single batch sample is processed at a time, significantly reducing peak memory usage. Under the same configuration, VRAM usage drops to approximately **0.83 GB**.

In this update, we extend the same idea further by introducing chunking across the **sequence dimension** as well. Instead of materializing logits for the entire `(batch_size × context_length)` space at once, we flatten these dimensions and process them in smaller chunks using a configurable multiplier. This allows Unsloth to support substantially longer contexts without increasing peak memory usage.

Equation 3: Logit Memory (GB)=context length×multiplier×vocab dim/1024^3

### 👻Hidden States Chunking

We also observed that at longer context lengths, hidden states can become a significant contributor to memory usage. For demonstration, we will assume `hidden_states_dim=4096`. 
Hidden States Memory (GB)=batch size×context length×hidden states dim/1024^3

With a `batch_size = 8` and `context_length = 64000`, this would result in a VRAM usage of approximately **2 GB**. In this release, we introduce optional chunking over the batch dimension for the hidden states tensor during log-probability computation. This would cause the VRAM usage to be divided by the batch size or in this case be **0.244 GB**.

This reduces the peak VRAM required to materialize hidden states.

### 🌵Offloading activations for log softmax

During the development of this release, we discovered that when tiling across the batch dimension for hidden states, the activations were not being offloaded after the fused logits and logprobs computation.  
To address this, we added explicit logic to offload these activations outside the model’s forward pass.

### ✨Configuring parameters:

If you do not configure `unsloth_grpo_mini_batch` and `unsloth_logit_chunk_multiplier`, we will **automatically tune these two parameters** for you based on your available VRAM and depending on the size of your context length. Below however is how you can change these variables in your GRPO run:

```python
training_args = GRPOConfig(
    ...
    unsloth_grpo_mini_batch = 3
    unsloth_logit_chunk_multiplier = 2
    ...
)
```

### 📼vLLM for RL

**For RL workflows, the inference/generation phase is the main bottleneck**. To address this, we utilize [vLLM](https://github.com/vllm-project/vllm), which has accelerated generation by up to 11x compared to normal generation.

Acknowledgements: A huge thank you to the Hugging Face team and libraries for powering Unsloth and making this possible.
