Back to News
Unsloth AIMay 13, 2026

Qwen3.6 GGUFs

Paid
LLM
GGUF
Local Inference
Qwen
Open Source
Multi-Token Prediction

Explore Qwen3.6 GGUFs

Visit the official website to learn more and get started

### TL;DR

Unsloth AI released experimental GGUFs for Qwen3.6 in 27B and 35B variants, featuring a new Multi-Token Prediction (MTP) architecture that delivers up to 220 tokens/s on a single GPU — a 1.4x speed-up over previous versions. The 27B model runs on 18GB RAM and the 35B-A3B on 22GB, making high-quality local inference more accessible.

Key Insights & Metrics

Pricing
Free (Open Source)
Cost structure
Version
3.6
Current release version
Hardware
18GB RAM for 27B; 22GB RAM for 35B-A3B
Compute requirements
Category
Paid
Licensing model
Region
Australia
Primary region

Key Features

  • MTP (Multi-Token Prediction) architecture for up to 220 tokens/s on a single GPU
  • 27B model runs on 18GB RAM; 35B-A3B runs on 22GB
  • 1.4x speed improvement over previous Qwen GGUF versions

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode