Hybrid Models
hybrid-1 is a gate model with a synthesizer that routes easy tasks to a cheap worker (Beta model) to save on total costs where subtasks are easy. Unlike a static model router, it works within the flow. Unlike a task difficulty classifier, it predicts whether Beta's completion actually matches what Alpha was going for. You can call it like any single model.
How it works
We recommend pairing models from the same family for Alpha and Beta (for example glm-5.3 with glm-5.3-flash) for the best results.
Usage
from openai import OpenAI
client = OpenAI(
base_url="https://api.thetokencompany.com/v1",
api_key="ttc-...",
)
response = client.chat.completions.create(
model="hybrid-1/glm-5.3+glm-5.3-flash",
messages=messages,
tools=tools,
)
Parallel runs
Each agent run is a session. Send an x-hybrid-session header with a unique ID per run, and the same ID on every request of that run. Without it, parallel runs still work independently, but a retried first request can't be told apart from a new run, so it runs and is billed again.
import uuid
run_id = str(uuid.uuid4()) # one per agent run
response = client.chat.completions.create(
model="hybrid-1/glm-5.3+glm-5.3-flash",
messages=messages,
tools=tools,
extra_headers={"x-hybrid-session": run_id},
)
Response
A standard chat completion, plus an x_hybrid field that shows which model answered and what the request cost.
{
"id": "chatcmpl-hyb3f9c2a61d0e8b4a7c5d19e2",
"object": "chat.completion",
"created": 1790718000,
"model": "glm-5.3-flash",
"choices": [...],
"usage": {...},
"x_hybrid": {
"session": "e3841d603ab7ac76",
"served_by": "beta",
"model": "glm-5.3-flash",
"cost": { "usd": 0.00041, "by_role": { "beta": 0.00041 } },
"version": "hybrid-1"
}
}Pricing
You pay list prices with no markup: input tokens for each model that reads your request, and output tokens only for the model whose answer you get. The gate and synthesis are free.
x_hybrid.cost.usd is the total for the request, and x_hybrid.cost.by_role shows how it splits between the Alpha and Beta models. When Alpha takes over from a Beta draft, both appear: Beta for reading your request, Alpha for reading it and writing the answer.