Ox Alpha processed two trillion tokens on OpenRouter in under two days.
Users still do not know who provides it. They also do not know whether their prompts and completions are being retained.
The Early Coding Result Looks Good, but Proves Little
The attraction is easy to understand. Ox Alpha is free, supports tool calling, and has a 1,048,576-token context window.
It also performed well in one early coding test:
- Ox Alpha: 80%
- Claude Fable 5: 65%
- GPT 5.6 Sol: 52%
Ox Alpha came out ahead on this 10-task Deep SWE run. But ten tasks cannot establish a reliable ranking.
Its stated uncertainty range was 49% to 94%. That is too wide to tell us precisely how strong the model is.
The result gives people a reason to test Ox Alpha. It does not prove that Ox Alpha is the best of these models.
Technical Clues Point Toward GLM
Ox Alpha may be anonymous, but it left technical clues.
Its tokenizer matched GLM 5.3 across six texts. Its video token counts matched GLM 5V Turbo exactly.
Those are meaningful connections to ZAI’s GLM family, but they do not confirm who operates Ox Alpha. Technical similarity cannot settle questions about the provider, its infrastructure, or its policies.
The Retention Claims Do Not Fit Together
More testing can reduce the benchmark uncertainty. The conflicting data-retention claims are more urgent.
OpenRouter warned that Ox Alpha’s provider retained prompts and completions.
OpenCode claimed:
“zero data retention.”
Both statements cannot provide users with a clear account of what happens to their data.
That matters when people may be submitting code, documents, tool inputs, and long conversations. After two trillion tokens, users still lack a confirmed answer about who handles that information and whether it is kept.
Ox Alpha may prove to be excellent, but strong benchmark results cannot answer who handles the data. Until the provider and retention policy are confirmed, do not send it sensitive information.