Hybrid models are only available to approved customers.

Hybrid Models

hybrid-1 is a gate model with a synthesizer that routes easy tasks to a cheap worker (Beta model) to save on total costs where subtasks are easy. Unlike a static model router, it works within the flow. Unlike a task difficulty classifier, it predicts whether Beta's completion actually matches what Alpha was going for. You can call it like any single model.

How it works

Alpha modelBeta model
Alpha model
Beta model
kimi-k3glm-5.3claude-sonnet-5.5gpt-6.1-solclaude-haiku-4.5

We recommend pairing models from the same family for Alpha and Beta (for example glm-5.3 with glm-5.3-flash) for the best results.

Usage

from openai import OpenAI

client = OpenAI(
base_url="https://api.thetokencompany.com/v1",
api_key="ttc-...",
)

response = client.chat.completions.create(
model="hybrid-1/glm-5.3+glm-5.3-flash",
messages=messages,
tools=tools,
)

Parallel runs

Each agent run is a session. Send an x-hybrid-session header with a unique ID per run, and the same ID on every request of that run. Without it, parallel runs still work independently, but a retried first request can't be told apart from a new run, so it runs and is billed again.

import uuid

run_id = str(uuid.uuid4()) # one per agent run

response = client.chat.completions.create(
model="hybrid-1/glm-5.3+glm-5.3-flash",
messages=messages,
tools=tools,
extra_headers={"x-hybrid-session": run_id},
)

Response

A standard chat completion, plus an x_hybrid field that shows which model answered and what the request cost.

{
"id": "chatcmpl-hyb3f9c2a61d0e8b4a7c5d19e2",
"object": "chat.completion",
"created": 1790718000,
"model": "glm-5.3-flash",
"choices": [...],
"usage": {...},
"x_hybrid": {
"session": "e3841d603ab7ac76",
"served_by": "beta",
"model": "glm-5.3-flash",
"cost": { "usd": 0.00041, "by_role": { "beta": 0.00041 } },
"version": "hybrid-1"
}
}

Pricing

You pay list prices with no markup: input tokens for each model that reads your request, and output tokens only for the model whose answer you get. The gate and synthesis are free.

x_hybrid.cost.usd is the total for the request, and x_hybrid.cost.by_role shows how it splits between the Alpha and Beta models. When Alpha takes over from a Beta draft, both appear: Beta for reading your request, Alpha for reading it and writing the answer.