Launching with Qwopus 3.6 27B v2 — Qwen 3.6 distilled on Claude Opus reasoning traces. Built for agentic reasoning (OpenClaw), capable coding (Pi, OpenCode, Claude Code subagents), and any OpenAI-compatible framework.
One email at launch with your discount code.
Your 20% discount code arrives on launch day.
Please try again or email hello@tokensunchained.com
The community builds extraordinary models — distills, fine-tunes, specialty coders. They live on HuggingFace with thousands of downloads and no production endpoint. We give them a home: clean APIs, proper infrastructure, friction-less integrations, and honest pricing.
Qwen 3.6 27B fine-tuned on Claude Opus 4.6/4.7 reasoning traces. Rather than copying Opus's compressed answer summaries, v2 uses Trace Inversion to reconstruct the full step-by-step reasoning — capturing how Opus decomposes tasks and calls tools.
Built on the Qwen 3.6 architecture (hybrid Gated DeltaNet + Attention, 262k native context). The v2 release posts measured gains over the base on reasoning quality and token efficiency — see the benchmarks.
Qwen 3.6 27B is an excellent open base. But like most reasoning models it spends heavily inside its <think> phase — often running long before it answers. Qwopus 3.6 27B v2's Trace-Inversion distillation targets exactly that: equal or better accuracy, reached with far fewer wasted tokens.
Creator-reported by Jackrong & Kyle Hessling on the v2 model card, measured on a 350-question MMLU-style set. Token and chain-length figures compare v2 to the Qwen 3.6 27B base on correctly-answered questions — accuracy goes up while token spend goes down. Independent PinchBench results will follow once our endpoint is live.
| Benchmark | Qwen 3.6 27B open base | Claude Opus 4.8 frontier |
|---|---|---|
| SWE-bench Verified | 77.2 | 88.6 |
| GPQA Diamond | 87.8 | 93.6 |
Base from the official Qwen 3.6 27B card (Qwen's harness); Opus 4.8 from Anthropic's system card, reported with max-effort adaptive thinking over multiple attempts — a different, more favorable harness. These are not directly comparable, and a 27B open model isn't at frontier parity. Shown only for scale: our case is a strong open model at a fraction of Opus's $5 / $25 pricing, not a parity claim.
OpenAI and Anthropic-compatible API. Works with OpenClaw, Pi, OpenCode, Claude Code, LangChain, and any raw HTTP client — no proxy needed.
Register at tokensunchained.com. Verify your email, then navigate to API Keys and create a key. It's shown only once — save it immediately.
Open ~/.openclaw/openclaw.json. Add the block provided in your dashboard to the models.providers array, then also add the model to agents.defaults.models — both entries are required or you'll see "model not allowed" errors. Run openclaw gateway config.apply --file ~/.openclaw/openclaw.json.
{
"name": "qwopus",
"baseUrl": "https://api.tokensunchained.com/v1",
"apiKey": "tu_your_key_here",
"api": "openai-completions",
"models": [{
"id": "qwopus-27b",
"name": "Qwopus 3.6 27B",
"reasoning": true,
"contextWindow": 262144
}]
}
"agents": { "defaults": { "models": { "qwopus/qwopus-27b": { "alias": "qwopus", "default": true } } } }
Paste the prompt provided in your dashboard into your current OpenClaw chat (replacing the API key). The agent will read its own config, apply both required changes, and restart the gateway. Review the diff before approving.
Please add TokensUnchained as a provider in my OpenClaw config.
1. Read ~/.openclaw/openclaw.json
2. Add this to models.providers[]:
{"name":"qwopus","baseUrl":"https://api.tokensunchained.com/v1",
"apiKey":"tu_your_key_here","api":"openai-completions",
"models":[{"id":"qwopus-27b","reasoning":true,"contextWindow":262144}]}
3. Also add to agents.defaults.models:
"qwopus/qwopus-27b":{"alias":"qwopus","default":true}
4. Show me the full diff before saving anything.
5. After I approve, save the file and run:
openclaw gateway config.apply --file ~/.openclaw/openclaw.json
Then run /models in OpenClaw to confirm qwopus-27b appears. Switch with /model qwopus and send a test message.
Add to your shell profile (~/.zshrc or ~/.bashrc), then restart your terminal.
export TU_API_KEY=tu_your_key_here
In Pi's LLM settings, add an OpenAI-compatible provider. Pi uses the standard OpenAI endpoint format — point it at our base URL with your key.
{
"provider": "openai-compatible",
"baseUrl": "https://api.tokensunchained.com/v1",
"apiKey": "tu_your_key_here",
"model": "qwopus-27b"
}
Send a multi-step problem. Look for the <thinking> block — structured decomposition before the answer is the Opus distillation in action.
OpenCode supports custom OpenAI-compatible providers via its configuration. Open the settings panel and add a new provider entry.
{
"providers": {
"tokensunchained": {
"name": "TokensUnchained",
"type": "openai",
"apiKey": "tu_your_key_here",
"baseURL": "https://api.tokensunchained.com/v1",
"models": ["qwopus-27b"]
}
}
}
Our API speaks both OpenAI and Anthropic formats natively. Point Claude Code at our /v1/messages endpoint — no proxy or gateway needed.
# Add to your shell profile (~/.zshrc or ~/.bashrc)
export ANTHROPIC_BASE_URL=https://api.tokensunchained.com
export ANTHROPIC_API_KEY=tu_your_key_here
export ANTHROPIC_DEFAULT_SONNET_MODEL=qwopus-27b
from openai import OpenAI client = OpenAI( api_key="tu_your_key_here", base_url="https://api.tokensunchained.com/v1" ) response = client.chat.completions.create( model="qwopus-27b", messages=[{"role": "user", "content": "Hello"}], stream=True )
curl https://api.tokensunchained.com/v1/chat/completions \ -H "Authorization: Bearer tu_your_key_here" \ -H "Content-Type: application/json" \ -d '{"model":"qwopus-27b","stream":true, "messages":[{"role":"user","content":"Hello"}]}'
The <think> block streams as a separate content field before the main response. Standard SSE format — no custom parsing needed in any OpenAI-compatible client.
Cancel anytime · Credits never expire
Blended rate at 7:1 input/output ratio. Opus 4.8 API is $5.00 in / $25.00 out per 1M ($7.50/1M blended).
Qwen 3.6 27B fine-tuned on Claude Opus 4.6/4.7 reasoning traces — the v2 release, built by developer Jackrong with Kyle Hessling. It uses a "Trace Inversion" method that reconstructs full step-by-step reasoning instead of copying the compressed answer summaries commercial models expose. The Qwen 3.6 base uses a hybrid Gated DeltaNet + Attention architecture with 262k native context and scores 77.2% on SWE-bench Verified and 87.8% on GPQA Diamond. On the creators' 350-question MMLU-Pro subset, v2 scores 87.43% vs the base's 84.86% — while using markedly fewer reasoning tokens to get there.
No — Opus is a frontier model; Qwopus is a 27B open distill, and we don't claim parity. What the distillation captures well is Opus's reasoning structure and tool-calling approach, which is what matters most for agent workflows. For scale: the Qwen 3.6 base scores 77.2% on SWE-bench Verified, versus Claude Opus 4.8 at 88.6% (on Anthropic's max-effort harness — not a like-for-like setup). The point isn't to match Opus; it's a capable open model at a fraction of Opus's $5/$25 pricing, with tighter, less wasteful reasoning than the base.
Pre-register before launch. On launch day we email you a single-use 20% discount code valid for your first subscription payment or your first PAYG credit top-up. No card required to pre-register.
Qwopus 3.6 27B v2 is released and is the model we're preparing to serve. Jackrong and Kyle Hessling continue to iterate — a 35B-A3B MoE variant is already out for high-throughput, thinking-off coding, which we're evaluating as a possible fast tier later. Right now the focus is the service: stable uptime, accurate billing, reliable streaming, and publishing our own independent benchmark runs once the endpoint is live. More models will follow as the community validates them.
One email at launch. Unsubscribe anytime.