Comparison
Ox Alpha vs Grok 4
Head-to-head on an independent benchmark of 10 real-world coding tasks: the stealth model against xAI's Grok 4.6, xhigh effort.
80%
Ox Alpha
ox-alpha · 8/10 tasks
VS
62%
Grok 4
grok-4.6 [xhigh] · mean pass rate
Task by task
Where each model wins.
4
tasks where Ox Alpha had the edge
2
tasks where Grok 4 had the edge
4
tasks both solved cleanly
0
tasks neither solved
| Task | grok-4.6 [xhigh] | ox-alpha | Verdict |
|---|---|---|---|
| anko-typed-variable-bindings | 1/4 | ✓ | Ox Alpha |
| arktype-json-schema-refs | 1/4 | ✓ | Ox Alpha |
| fastapi-deprecation-headers | 4/4 | ✓ | Both solved |
| helm-unified-manifest-stream | 4/4 | ✓ | Both solved |
| igel-persist-feature-schema | 4/4 | ✓ | Both solved |
| katex-multicolumn-array-spans | 4/4 | ✓ | Both solved |
| meriyah-explicit-resource-decl | 0/4 | ✓ | Ox Alpha |
| query-persist-restored-state | 2/4 | ✓ | Ox Alpha |
| scc-bounded-memory-spilling | 4/4 | ✕ | Grok 4 |
| vulture-persistent-analysis-cache | 1/4 | ✕ | Grok 4 |
| Mean on these 10 | 62% | 80% |
Independent community benchmark — 10 real-world coding tasks. Reference models were scored as passes out of 4 attempts per task; Ox Alpha was recorded as pass/fail. "Edge" = Ox Alpha passed while Grok 4 missed at least one attempt, or vice-versa. Third-party data presented as published — 10 tasks is a small sample, so read it as directional.
What the numbers say
Ox Alpha vs Grok 4.6, xhigh effort
- 8 of 10 solved by Ox Alpha, versus 5 tasks that Grok 4 passed on all four attempts.
- Solved where Grok 4 went 0/4: meriyah-explicit-resource-decl.
- Grok 4's strongest ground: 2 tasks it passed that Ox Alpha missed — scc-bounded-memory-spilling, vulture-persistent-analysis-cache.
- Both stumbled on 0 tasks — a reminder that no model is a guaranteed pass on real-world bugs.
Ox Alpha at a glance
What's known about the stealth model
- Type
- Reasoning model
- Context window
- 1,048,576 tokens
- Max output
- 131,072 tokens
- Inputs
- Text · Image · Video
- Developer features
- Tool calling · JSON output
- Creator
- Unknown (stealth)
About Grok 4
xAI · benchmarked build: grok-4.6 [xhigh]
- xAI's Grok 4.6, run at "xhigh" reasoning effort.
- Mean pass rate on these 10 tasks: 62% — passing at least one attempt on 9 of 10.
Why it matters
Reading a 10-task benchmark honestly
- These are real repository tasks, not toy puzzles — the kind of work a reasoning model is built for.
- Ten tasks is a small sample; a few tasks either way would move the means. Use it to decide what to try, not to declare a winner.
- The best benchmark is your own problem. Ox Alpha is free to try right now.
Run your own benchmark.
Paste the bug, the spec, or the whole repo. Ox Alpha is free while it's in stealth.
Try Ox Alpha