Coming Soon

Bleeding-edge open-weight models
served on a platter

Launching with Qwopus 3.6 27B v2 — Qwen 3.6 distilled on Claude Opus reasoning traces. Built for agentic reasoning (OpenClaw), capable coding (Pi, OpenCode, Claude Code subagents), and any OpenAI-compatible framework.

From $29/mo · PAYG $3.00 in / $7.00 out per 1M
🎁
Launch offer — 20% off your first purchase
Pre-register before launch and get a 20% discount code for your first subscription or credit top-up.
v2 released · API launching soon
Qwopus 3.6 27B v2 · SGLang · 262k context OpenAI + Anthropic API compatible Opus-style reasoning No hourly windows Zero data retention

Bringing the excellence of open-weight models to everyone.

The community builds extraordinary models — distills, fine-tunes, specialty coders. They live on HuggingFace with thousands of downloads and no production endpoint. We give them a home: clean APIs, proper infrastructure, friction-less integrations, and honest pricing.

Flagship model
Qwopus 3.6 27B v2

Qwen 3.6 27B fine-tuned on Claude Opus 4.6/4.7 reasoning traces. Rather than copying Opus's compressed answer summaries, v2 uses Trace Inversion to reconstruct the full step-by-step reasoning — capturing how Opus decomposes tasks and calls tools.

Built on the Qwen 3.6 architecture (hybrid Gated DeltaNet + Attention, 262k native context). The v2 release posts measured gains over the base on reasoning quality and token efficiency — see the benchmarks.

Agent orchestration Coding Vision

Strong base, tighter reasoning

Qwen 3.6 27B is an excellent open base. But like most reasoning models it spends heavily inside its <think> phase — often running long before it answers. Qwopus 3.6 27B v2's Trace-Inversion distillation targets exactly that: equal or better accuracy, reached with far fewer wasted tokens.

v2 vs the Qwen 3.6 27B base it's built on
−36%
Tokens to a
correct answer
−52%
Reasoning
chain length
+2.57 pp
MMLU-Pro subset
87.43% vs 84.86%
1.66×
MTP decode
speedup

Creator-reported by Jackrong & Kyle Hessling on the v2 model card, measured on a 350-question MMLU-style set. Token and chain-length figures compare v2 to the Qwen 3.6 27B base on correctly-answered questions — accuracy goes up while token spend goes down. Independent PinchBench results will follow once our endpoint is live.

For scale — where the open base sits vs the frontier
Benchmark Qwen 3.6 27B open base Claude Opus 4.8 frontier
SWE-bench Verified77.288.6
GPQA Diamond87.893.6

Base from the official Qwen 3.6 27B card (Qwen's harness); Opus 4.8 from Anthropic's system card, reported with max-effort adaptive thinking over multiple attempts — a different, more favorable harness. These are not directly comparable, and a 27B open model isn't at frontier parity. Shown only for scale: our case is a strong open model at a fraction of Opus's $5 / $25 pricing, not a parity claim.

Running in 2 minutes

OpenAI and Anthropic-compatible API. Works with OpenClaw, Pi, OpenCode, Claude Code, LangChain, and any raw HTTP client — no proxy needed.

01

Sign up and get your API key

Register at tokensunchained.com. Verify your email, then navigate to API Keys and create a key. It's shown only once — save it immediately.

02a

Manual — edit openclaw.json directly

▼

Open ~/.openclaw/openclaw.json. Add the block provided in your dashboard to the models.providers array, then also add the model to agents.defaults.models — both entries are required or you'll see "model not allowed" errors. Run openclaw gateway config.apply --file ~/.openclaw/openclaw.json.

Add to models.providers[ ] in openclaw.json
{
  "name": "qwopus",
  "baseUrl": "https://api.tokensunchained.com/v1",
  "apiKey": "tu_your_key_here",
  "api": "openai-completions",
  "models": [{
    "id": "qwopus-27b",
    "name": "Qwopus 3.6 27B",
    "reasoning": true,
    "contextWindow": 262144
  }]
}
Also add to agents.defaults.models (required)
"agents": {
  "defaults": {
    "models": {
      "qwopus/qwopus-27b": {
        "alias": "qwopus",
        "default": true
      }
    }
  }
}
02b

Auto — let your current agent do it

▼

Paste the prompt provided in your dashboard into your current OpenClaw chat (replacing the API key). The agent will read its own config, apply both required changes, and restart the gateway. Review the diff before approving.

Paste into your OpenClaw chat
Please add TokensUnchained as a provider in my OpenClaw config.

1. Read ~/.openclaw/openclaw.json
2. Add this to models.providers[]:
   {"name":"qwopus","baseUrl":"https://api.tokensunchained.com/v1",
    "apiKey":"tu_your_key_here","api":"openai-completions",
    "models":[{"id":"qwopus-27b","reasoning":true,"contextWindow":262144}]}
3. Also add to agents.defaults.models:
   "qwopus/qwopus-27b":{"alias":"qwopus","default":true}
4. Show me the full diff before saving anything.
5. After I approve, save the file and run:
   openclaw gateway config.apply --file ~/.openclaw/openclaw.json
03

Apply and verify

Then run /models in OpenClaw to confirm qwopus-27b appears. Switch with /model qwopus and send a test message.

01

Get your API key and set environment variable

Add to your shell profile (~/.zshrc or ~/.bashrc), then restart your terminal.

Shell environment
export TU_API_KEY=tu_your_key_here
02

Configure Pi to use the custom provider

In Pi's LLM settings, add an OpenAI-compatible provider. Pi uses the standard OpenAI endpoint format — point it at our base URL with your key.

Pi custom provider config (via Pi settings UI or config file)
{
  "provider": "openai-compatible",
  "baseUrl": "https://api.tokensunchained.com/v1",
  "apiKey": "tu_your_key_here",
  "model": "qwopus-27b"
}
03

Test reasoning output

Send a multi-step problem. Look for the <thinking> block — structured decomposition before the answer is the Opus distillation in action.

01a

Add a custom provider in OpenCode settings

OpenCode supports custom OpenAI-compatible providers via its configuration. Open the settings panel and add a new provider entry.

OpenCode provider config (~/.config/opencode/config.json or settings UI)
{
  "providers": {
    "tokensunchained": {
      "name": "TokensUnchained",
      "type": "openai",
      "apiKey": "tu_your_key_here",
      "baseURL": "https://api.tokensunchained.com/v1",
      "models": ["qwopus-27b"]
    }
  }
}
01b

Claude Code — direct connection

Our API speaks both OpenAI and Anthropic formats natively. Point Claude Code at our /v1/messages endpoint — no proxy or gateway needed.

Launch Claude Code with Qwopus
# Add to your shell profile (~/.zshrc or ~/.bashrc)
export ANTHROPIC_BASE_URL=https://api.tokensunchained.com
export ANTHROPIC_API_KEY=tu_your_key_here
export ANTHROPIC_DEFAULT_SONNET_MODEL=qwopus-27b
Python — openai library
from openai import OpenAI

client = OpenAI(
    api_key="tu_your_key_here",
    base_url="https://api.tokensunchained.com/v1"
)

response = client.chat.completions.create(
    model="qwopus-27b",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True
)
cURL
curl https://api.tokensunchained.com/v1/chat/completions \
  -H "Authorization: Bearer tu_your_key_here" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwopus-27b","stream":true,
     "messages":[{"role":"user","content":"Hello"}]}'

Reasoning tokens in streams

The <think> block streams as a separate content field before the main response. Standard SSE format — no custom parsing needed in any OpenAI-compatible client.

Pricing

Power User

$29 /mo
Billed monthly
10M tokens / mo
333k tokens / day
Prompts
~165 / day
Chat, Q&A, short code
Agent tasks
~22 / day
Refactors, multi-file edits
17% cheaper than PAYG. No hourly windows.
  • Covers a full day of coding
  • Overage: $3.00/$7.00 per 1M

Unchained

$119 /mo
Billed monthly
55M tokens / mo
1.83M tokens / day
Prompts
~915 / day
Chat, Q&A, short code
Agent tasks
~122 / day
Refactors, multi-file edits
38% cheaper than PAYG. TokensUnchained
  • Heavy multi-agent orchestration
  • Overage: $2.00/$5.00 per 1M
Prompt ≈ 2k tokens (short chat exchange) · Agent task ≈ 15k tokens (codebase refactor, multi-file edit, complex reasoning chain). Actual usage varies by context length and output size.
Get 20% Off Your First Purchase

Cancel anytime · Credits never expire

$3.00 in / $7.00 out per 1M tokens


Our PAYG rate
$3.50/1M blended
vs Opus 4.8 API
2.1× cheaper
Subscribers save up to
38% off PAYG

Blended rate at 7:1 input/output ratio. Opus 4.8 API is $5.00 in / $25.00 out per 1M ($7.50/1M blended).

Common questions

What is Qwopus 3.6 27B?

Qwen 3.6 27B fine-tuned on Claude Opus 4.6/4.7 reasoning traces — the v2 release, built by developer Jackrong with Kyle Hessling. It uses a "Trace Inversion" method that reconstructs full step-by-step reasoning instead of copying the compressed answer summaries commercial models expose. The Qwen 3.6 base uses a hybrid Gated DeltaNet + Attention architecture with 262k native context and scores 77.2% on SWE-bench Verified and 87.8% on GPQA Diamond. On the creators' 350-question MMLU-Pro subset, v2 scores 87.43% vs the base's 84.86% — while using markedly fewer reasoning tokens to get there.

Is it as good as Claude Opus?

No — Opus is a frontier model; Qwopus is a 27B open distill, and we don't claim parity. What the distillation captures well is Opus's reasoning structure and tool-calling approach, which is what matters most for agent workflows. For scale: the Qwen 3.6 base scores 77.2% on SWE-bench Verified, versus Claude Opus 4.8 at 88.6% (on Anthropic's max-effort harness — not a like-for-like setup). The point isn't to match Opus; it's a capable open model at a fraction of Opus's $5/$25 pricing, with tighter, less wasteful reasoning than the base.

How does the 20% launch discount work?

Pre-register before launch. On launch day we email you a single-use 20% discount code valid for your first subscription payment or your first PAYG credit top-up. No card required to pre-register.

What's coming after launch?

Qwopus 3.6 27B v2 is released and is the model we're preparing to serve. Jackrong and Kyle Hessling continue to iterate — a 35B-A3B MoE variant is already out for high-throughput, thinking-off coding, which we're evaluating as a possible fast tier later. Right now the focus is the service: stable uptime, accurate billing, reliable streaming, and publishing our own independent benchmark runs once the endpoint is live. More models will follow as the community validates them.

Pre-register
20% off your first purchase.