OpenRouter's Next Free Stealth Model Expires in Three Days.
stealth/space-bunny-alpha went live on OpenRouter September 23, costs nothing, claims a 1-million-token context window, and expires October 5 — three days from this post. So I ran the same scripted battery I used on ox-alpha and union-alpha — three separate full runs, not one, so there's an actual rate to report instead of a single data point — pushed the context claim past its own advertised ceiling to see where it actually breaks, and ran it head-to-head against five named models on the same gateway, same day. Short version: the context number holds up almost exactly, and the failure past it is a real error instead of a fake one. Two contract parameters are quietly ignored, and 3 of 48 calls across those three sessions came back HTTP 200 while actually failing — all three in one session, zero in the other two. On identity: this is MiniMax at 97% confidence — a reproduced 36-for-36 tokenizer match, not a guess. Specifically M3.1-Flash-Preview is a separate, weaker claim at 70–75%, built on three matching specs rather than a tokenizer file I don't have. My first pass guessed Llama and was wrong before any of that.
| Claim | Verdict | Confidence / evidence |
|---|---|---|
| Context window (1M advertised) | Holds | Clean to 974,950 tokens, honest 400 past 1,000,000 — the real and advertised ceilings match |
| Parallel tool calls | Clean | 3/3 sessions, no duplicate-key bug |
Strict JSON schema (response_format) | Broken — and it's MiniMax, not the route | 0/9 runs across 3 sessions; reproduces identically on named minimax-m3:cloud |
Forced tool_choice | Ignored | 0/3 sessions obey it |
| Disguised HTTP-200 failures | Real, bursty | 3/48 calls across 3 sessions, clustered in one |
reasoning_tokens billing field | Broken | Reports 0 in 9/9 checks despite real, visible reasoning |
| Identity: MiniMax family | Proven | 97% — 36/36 raw-tokenizer pairwise match, reproduced |
| Identity: specifically M3.1-Flash-Preview | Likely, not proven | 70–75% — 3 spec matches, no tokenizer file for that exact checkpoint |
Three days, then it's gone
This is the third free stealth model I've run through this process. Ox-alpha in August, union-alpha in September, and now space-bunny-alpha, pulled live from OpenRouter's registry this morning:
created: 2026-09-23
expiration_date: 2026-10-05
pricing: { "prompt": "0", "completion": "0" }
context_length: 1,000,000
tokenizer: "Other"
Twelve-day preview window. Three days left as I write this. The last two posts in this series could argue "architect so the free price doesn't matter later." This one doesn't get that luxury — it's either still around in some form after October 5 or it isn't. Either way, the only useful thing to do with three days is find out what's actually true right now.
Don't guess — inspect.
Especially with three days on the clock.
What the registry claims
| Claim | Value from the registry |
|---|---|
| Context window | 1,000,000 tokens |
| Max output | 524,288 tokens |
| Modalities | text+image+video → text |
| Tool calling / structured output | both listed as supported |
| Reasoning | mandatory — five effort tiers: low / medium / high / xhigh / max, default max |
| Pricing | $0 input, $0 output |
| Tokenizer | listed as "Other" — no vendor attribution |
| Expires | 2026-10-05 |
Two things stand out before a single test runs. The effort dial has five rungs instead of ox-alpha's three, and reasoning is marked mandatory with a default of max — the most aggressive default I've seen advertised on this gateway. And the tokenizer field tells you nothing, same as every stealth listing. More on that later.
How I tested
Same discipline as the last two rounds, rebuilt from scratch for this model: every ground truth is computed in the test script, not memorized. Expected products, expected letter counts, expected vault codes planted at known depths in a haystack the script generated, expected shapes and serial numbers baked into a PNG the script drew. The grader is regex and arithmetic.
Before running anything at scale, I validated one call end-to-end — same rule I apply to every paid API, doubled for a free one with an expiration date, because a free meter is exactly the kind of thing nobody reads errors off of.
The 16-call core battery (reasoning, schema, tools, format, effort dial) ran three separate times, not once. A single run tells you what happened; it doesn't tell you what's reliable. Three runs are what turned "4 disguised failures out of 48 calls" into "3 failures, all in one session, zero in the other two" — a materially different and more useful finding than the single-session number would have been.
- Objective reasoning — multiplication, a six-step chained word problem, case-insensitive letter counting, a bat-and-ball trap with changed numbers.
- Structured output — strict JSON Schema extraction, three runs.
- Tool calling — single call, parallel calls across two cities plus a calculation, and a forced
tool_choice. - Long context — a needle-in-haystack ladder run from 100k tokens up past the advertised 1,000,000-token ceiling, with a fourth, never-planted needle as an honesty probe.
- Vision — a rendered image with known text and four known shapes.
- Head-to-head — the structural tests replayed verbatim against five models I verified were still live on OpenRouter today, not models I remembered being live two months ago.
- Tokenizer fingerprint — same-gateway prompt-token residuals across nine adversarial strings against five named peers, plus raw pairwise comparison against four local
tokenizer.jsonfiles.
What held up
| Test | Result | Detail |
|---|---|---|
| Reasoning battery | 4/4 | 335,412; the 86.95 decimal chain; letter count 12; bat/ball = 10¢. All correct. |
| Single & parallel tool calls | clean | Three separate, legally-formed calls for Tokyo, Paris, and 12*9. No duplicate-key collapse — the exact bug that broke ox-alpha's agent loop didn't reproduce here. |
| Exact-format instruction following | clean | Five-line LINE1..LINE5 contract honored exactly. |
| Vision | 5/5 | Large code, small-print serial, all four shapes with correct colors and positions. 3.7 seconds wall time. |
| Long-context retrieval | clean to 974,950 reported tokens | All three needles found at every rung from 100k up to just shy of the advertised ceiling. |
| Honesty at depth | perfect | The never-planted fourth needle came back NONE at every successful rung — no confabulation anywhere in the haystack. |
| Context ceiling behavior | honest | Pushed to ~1,127,000 tokens: a clean HTTP 400 naming the exact limit and the exact overage. No silent empty success. |
| Streaming speed | fast | 1.03s to first token, ~116 tokens/sec sustained on a ~250-word technical paragraph. |
| Cold start | fast | First call of the session: 1.3 seconds. Ox-alpha's first call took 87. |
The long-context result matters most here. I didn't stop at the advertised number — I kept pushing the haystack bigger until the API refused it. Retrieval held clean, including the honesty probe, all the way to 974,950 reported prompt tokens. The next rung up (~1,127,000 tokens) got a real HTTP 400 naming the exact overage:
"This endpoint's maximum context length is 1000000 tokens. However, you
requested about 1129425 tokens (1127425 of text input, 2000 in the output).
Please reduce the length of either one, or use the context-compression
plugin to compress your prompt automatically."
Compare that to ox-alpha: past ~660k tokens, ox-alpha returned HTTP 200, finish_reason: "stop", and zero bytes of content. A fake success. Space-bunny-alpha does the opposite at its own real ceiling — a specific error that tells you exactly what to change. Advertised limit and actual limit are the same number, and the failure past that number looks like a failure. That's the whole deal with trusting a context window.
Where it broke
1. Strict JSON schema is a no-op, not a crash
The registry lists response_format as supported. I sent a strict JSON Schema asking for a specific shape — order_id, customer.tier, an items array with sku/qty/unit_price_usd, a total_usd. Three runs. Here's run 1, verbatim and unedited:
{
"order_id": "ORD-7712",
"customer": {
"name": "Mara Quill",
"type": "enterprise"
},
"line_items": [
{ "sku": "SKU-A19", "quantity": 3, "unit_price": 12.50, "currency": "USD" },
{ "sku": "SKU-B22", "quantity": 1, "unit_price": 199.99, "currency": "USD" }
],
"total": 237.49,
"currency": "USD",
"priority": "high"
}
Notice what this is not: not markdown-fenced, not an apology about a missing schema, not garbled. It's clean, valid JSON. It's just the wrong shape — customer.type instead of customer.tier, line_items/quantity/unit_price instead of items/qty/unit_price_usd, no total_usd key at all. The content comprehension is actually correct every time — right customer, right items, right quantities, right arithmetic (237.49, computed correctly all three runs) — the model just extracts it as if you'd asked in plain prose and never saw your schema. I ran this three-attempt test in all three sessions: 9 runs total, 9 different field-naming choices, zero that match the contract. One session even came back markdown-fenced on its first attempt, where this session didn't — the exact presentation varies, the failure to honor the schema never does. response_format is listed as a supported parameter on this route and functions as a no-op.
Then I checked whether that's a stealth-route problem or a MiniMax problem, by running the identical extraction against the real, named minimax-m3:cloud on Ollama Cloud — official distribution, same weights, no stealth wrapper anywhere in the path. Same result: customer.type instead of tier, quantity/unit_price instead of qty/unit_price_usd, markdown-fenced this time, plus a paragraph of unsolicited commentary ("If your target schema differs... let me know"). That settles it: this is a MiniMax-M3 trait, not something OpenRouter's stealth gateway is doing to the request. Worth the contrast with the head-to-head table below — GLM, Grok, DeepSeek, and Claude Opus all honored this exact schema cleanly on the first try. Not every model struggles with this. This one does, with or without the stealth badge.
2. Forced tool_choice is ignored
I forced tool_choice: {"type": "function", "function": {"name": "calculator"}} on the message "Just say hello to me. Do not check anything." The model's full response:
content: "Hello! 👋"
tool_calls: null
Correct social read of the room. Wrong contract. The parameter said you must call this function; nothing enforced it, and no error surfaced either — the call just quietly did the socially-sensible thing instead of the specified thing. Same result in all three sessions: 3 for 3, ignored every time.
3. The real bug isn't downtime. It's lying about the downtime.
To be precise about what's being reported here: an anonymous stealth backend occasionally being unavailable is not a bug. It's the expected cost of running on something with less redundancy than a named flagship model, and I'd say the same about any provider. The bug is what happens when it is unavailable — the gateway answers with HTTP 200, the status code that means "here is your successful response," instead of a 502 or 503 that means "something failed, try again." That's a protocol-honesty problem, not an uptime complaint.
This is the one that matters most if you're building anything on this route. A single session isn't a valid sample for a failure rate, so I ran the same 16-call core battery three separate times instead of trusting one run. Combined: 3 of 48 calls came back as HTTP 200 with a body that is not a chat completion at all — but all three landed in the same session, and the other two sessions had zero. Here's one of them:
HTTP 200 OK
{
"id": "gen-1790951418-3fl2cX8KEseF2cHYx2eq",
"error": {
"message": "Provider returned an empty response",
"code": 502,
"metadata": { "error_type": "provider_unavailable" }
}
}
There is no choices array. There is no content. There is a 502-flavored error nested inside a 200 status code. Any pipeline — and this covers most pipelines, including my own first pass at the test client before I caught it — that checks only the HTTP status code and reaches for response.choices[0].message.content will either throw on a missing key or, worse, catch that and move on as if the call simply returned nothing. It is not rate limiting (no Retry-After, no 429). A clean retry of the exact same call always succeeded immediately after. And the gateway isn't incapable of honest error codes — the context-ceiling test earlier in this post got a correctly-coded 400 the moment the request actually exceeded the limit. Same route, same session, two different failure modes: one reported honestly, one not.
| Session | Core-battery calls | Disguised-200 failures |
|---|---|---|
| 1 | 16 | 3 |
| 2 | 16 | 0 |
| 3 | 16 | 0 |
| Combined | 48 | 3 (6.2%) |
That changes the conclusion in a useful way: this isn't a steady ~8% background failure rate, it's bursty — something went wrong for a stretch of one session and then stopped. Three calls out of forty-eight is still three calls too many if your pipeline trusts the status code alone, but "bursty and clustered" points at a transient upstream issue (a provider blip, a capacity event) rather than a structural one-in-twelve tax on every call. Either way, the fix is the same: check for choices, not just status code.
4. The reasoning-effort dial is real, but its own billing field lies about it
The registry's "mandatory reasoning, five effort tiers" claim checks out — sometimes. On "just say hello," no reasoning trace appeared at all, mandatory or not. On a harder arithmetic prompt, the model visibly reasons — the raw reasoning field returns real, legible chain-of-thought:
"We need answer only integer. Need compute 6839*5724 accurately. Let's
calculate. 5724*(6800+39) = 5724*6800 + 5724*39. 5724*68*100: 5724*68 =
*(70-2) = 400680-11448 = 389232; *100 = 38,923,200. ..."
I ran this same effort-dial check across all three sessions rather than trusting one:
| Session | low | high | max | reasoning_tokens (all 3 levels) |
|---|---|---|---|---|
| 1 | 58 | 105 | 171 | 0, 0, 0 |
| 2 | 81 | 114 | 180 | 0, 0, 0 |
| 3 | 96 | 88 | 127 | 0, 0, 0 |
| Average | 78 | 102 | 159 | — |
Session 3 isn't strictly monotonic — high cost fewer tokens than low that run. The averages still trend the right way (78 → 102 → 159), so the dial does something real on average, but don't expect every single call to respect the ordering; this looks like normal sampling variance in how much the model decides to think, not a broken dial. What's consistent across all three sessions, with zero exceptions in nine checks: usage.completion_tokens_details.reasoning_tokens reported exactly 0 every time, including on calls where the visible reasoning text ran to several hundred characters. If you're doing capacity planning or cost attribution off that field — which is the field OpenAI-compatible tooling is built to read for exactly this purpose — you will undercount this model's actual thinking cost by the full amount every time. Budget off completion_tokens, not the reasoning-specific sub-field, on this route.
Hidden-tax footnote. Every call, regardless of content, carries a fixed ~149-token prefix reported as cached_tokens — rising to 181 the moment you attach tool definitions (the tools array itself adds 32 tokens to that fixed, cached prefix). This held constant across 25+ varied calls in this session: trivial prompts, vision calls, long-context calls. There's a fixed system preamble on every conversation before your prompt even starts, and it's big enough that a one-word request like "PONG" spends more of its prompt budget on that prefix than on your actual message.
Then I made it fight the neighbors — with today's roster, not August's
Model slugs on this gateway don't stay still for two months. Before running the head-to-head, I pulled the live model list rather than reuse names from the ox-alpha or union-alpha posts — claude-opus-5 and glm-5.3 are both gone from the catalog already, replaced by claude-opus-4.7 and glm-5. I ran the same three structural prompts, verbatim, against five models confirmed live today:
| Model | Strict schema | Parallel tools | Forced choice | Notes |
|---|---|---|---|---|
| stealth/space-bunny-alpha | ✗ wrong shape | ✓ | ✗ ignored | free (preview, 3 days left) |
| z-ai/glm-5 | ✓ | ✓ | ✓ | reasons by default (607/46/55 reasoning tokens) |
| x-ai/grok-4.7 | ✓ | ✓ | ✓ | reasons by default (143/35/90 reasoning tokens) |
| deepseek/deepseek-v3.2-exp | ✓ | ✓ | ✓ | cheapest run of the five, $0.0004 total |
| anthropic/claude-opus-4.7 | ✓ | ✓ | ✓ | — |
| anthropic/claude-fable-5.1 | ✓ | ✓ | n/a — HTTP 400 | this Claude route rejects forced tool_choice outright rather than ignoring it |
Total spend across all five paid comparisons: about six cents. Two things fall out of this table. First, the GLM failure signature from my ox-alpha post in August is simply gone — glm-5 passes every test here cleanly. Six weeks was enough time for someone to fix that plumbing, which means failure-mode matching is a much weaker identity signal now than it was in August, and I'm not going to lean on it the way that post did. Second, every live model I could still name on this gateway today handles at least the strict-schema and parallel-tool contracts that space-bunny-alpha doesn't. That rules out "all stealth-tier routes on OpenRouter are like this" as an excuse. This is specific to this route, not a tax on anonymity in general.
So what is it? First the wrong lead, then the real one
The union-alpha post taught me not to do this: an earlier draft of it inherited the previous model's identity guess before the real battery ran, and the data contradicted it. So here's the full trail this time, including the part where I got it wrong first.
Solid regardless of identity: the hidden-preamble tax. The ~149-cached-token prefix on every call (181 with tools attached) isn't a guess. It's a flat, repeated measurement across 25+ calls this session, and it reproduced again on every re-check below.
My first identity pass, and why it was weak. I ran the same nine adversarial strings same-gateway against five named peers — Llama-3.1-8B, Llama-3.3-70B, GLM-4.7-Flash, Qwen3.5-27B, DeepSeek-V3.2-exp — and compared the residual (candidate token count minus space-bunny-alpha token count) string by string. GLM and Qwen diverged hard (std. dev. 19.0 and 41.6) because both have digit-efficient BPE vocabularies that space-bunny-alpha's tokenizer doesn't share. DeepSeek and the two Llama variants were flatter (std. dev. 4.1–5.1), which is real signal against GLM/Qwen but nowhere near proof of anything specific. I called that a 55–60%-confidence lean toward "Llama-family-shaped." That published number was honest about its own weakness, but it also wasn't the right family — because I hadn't thought to test the one lead that actually mattered yet.
Going back for a second pass. While researching this model after a first draft, I found one writeup, at cellcog.ai, covering the same question. It cited a tokenizer test from a group it calls "MarMar Labs" hitting a clean match across its own probe set, and named MiniMax's M3.1-Flash-Preview as the likely base. That's the one outside lead I had — one post, found by digging around, not a chorus of community chatter. I hadn't tested MiniMax at all in my own first pass — an obvious gap in hindsight, since MiniMax is a live, named family on the same OpenRouter gateway I was already probing everything else on. So instead of taking cellcog.ai's word for it, I went and ran it myself: same methodology, same nine probes, same-gateway, against every MiniMax model OpenRouter actually lists — minimax-m3, minimax-m2.7, minimax-m1, and minimax-01. (MiniMax's "M3.1-Flash-Preview" is not a live slug on OpenRouter's catalog as of this session, so I could not test that exact point release directly through a chat call — more on what that does and doesn't mean below.)
| Candidate | Residual vs space-bunny-alpha | Std. dev. (9 probes) | Reproduced? |
|---|---|---|---|
| minimax/minimax-m1 | exactly +290, every probe | 0.00 | yes — identical on independent re-run |
| minimax/minimax-01 | exactly +555, every probe | 0.00 | yes — identical on independent re-run |
| minimax/minimax-m2.7 | −113 to −115 | 0.83 | close but not exact — one probe shifted by 2 tokens on re-run |
| minimax/minimax-m3 (OpenRouter-served) | +7 to +20 (bimodal) | 6.13 | noisy at this endpoint — resolved below with the raw tokenizer file |
minimax-m1 and minimax-01 already land on a perfectly flat offset across nine structurally different strings, which doesn't happen between unrelated vocabularies by chance. But minimax-m3 — the one actually named after the generation cellcog.ai's source pointed at — came back noisy at the API layer (std. dev. 6.13), and I wasn't going to round that up to a match just because it was the closest-sounding name. So I went one level deeper than any API usage field: I pulled MiniMaxAI/MiniMax-M3's real, published tokenizer.json straight from Hugging Face and ran the same nine probes through it directly, no chat endpoint, no routing wrapper, nothing but raw tokenizer.encode() calls — the identical method I'd already used for the Llama/Mistral/Qwen local-file comparisons earlier in this post.
Once that worked, I didn't stop at one checkpoint. I pulled the published tokenizer.json for every MiniMax generation on Hugging Face with one — M1, M2, M2.1, M2.5, M2.7, and M3 — and ran the same raw test against all six. The files themselves aren't identical: M1 has its own file hash, M3 has its own, and all four M2.x releases in between share one identical file. But every single one of the six, hash differences and all, produces the exact same flat −156 offset against space-bunny-alpha, stdev 0.0000. MiniMax has kept the same effective vocabulary, at least for everything these nine probes touch, across its entire public lineage.
| Probe | MiniMax-M3 raw tokenizer | space-bunny-alpha (usage) | Delta |
|---|---|---|---|
| en_pangram | 10 | 166 | −156 |
| emoji_zwj | 21 | 177 | −156 |
| zh_long | 20 | 176 | −156 |
| fr | 17 | 173 | −156 |
| ru | 19 | 175 | −156 |
| ja | 10 | 166 | −156 |
| digits | 67 | 223 | −156 |
| code | 30 | 186 | −156 |
| repeat_eq | 3 | 159 | −156 |
Every probe, the same delta: exactly −156. Standard deviation 0.00. Pairwise cross-check (comparing every probe against every other probe, which cancels any constant wrapper and tests the raw shape of the vocabulary itself): 36 of 36 exact matches. That's not a lean and it's not a hint — a 36-for-36 pairwise hit across nine content types this different, from a vocabulary pulled straight from the model's own public weights, is the same bar I set for a fingerprint "hit" back in the ox-alpha post, and this clears it completely. The earlier noise at the OpenRouter-served minimax-m3 endpoint turns out to be exactly that — noise from that endpoint's own serving-layer wrapper, not the underlying vocabulary. The raw tokenizer file was the right test all along; the API usage field was a noisier proxy for it.
That resolves the tokenizer half of what cellcog.ai's writeup claimed — the "MarMar Labs" 36-for-36 match against MiniMax's M3 tokenizer, which I just reproduced myself with a primary-source test instead of taking it on their say-so. The other half of their case was a spec comparison naming M3.1-Flash-Preview specifically, and I checked that independently too rather than repeating it on their word: MiniMax's own M3.1-Flash-Preview, announced September 27, 2026, carries a 1,000,000-token context window and the identical five-tier reasoning-effort ladder — low, medium, high, xhigh, max — that space-bunny-alpha's OpenRouter listing advertises, confirmed across three independent trade write-ups of that announcement. Public reporting also notes M3.1-Flash-Preview normally runs only inside MiniMax's own coding tool rather than through a public API slug — which is exactly why I couldn't find or test that exact checkpoint name on OpenRouter directly, and fits neatly with "an anonymous stealth route is how you quietly put a gated model in front of outside traffic."
The tokenizer test isn't the only cross-check available. Since the lead is specifically MiniMax, I ran the same reasoning battery and vision test against the real, named minimax-m3:cloud on Ollama Cloud — official distribution, no stealth badge anywhere.
| Test | stealth/space-bunny-alpha | minimax-m3:cloud (named) |
|---|---|---|
| Reasoning battery | 4/4 | 4/4 |
| Vision (same test image) | 5/5 | 5/5 |
| Parallel tool calls | Clean | Clean |
| Strict JSON schema | Broken — wrong field names | Broken — same wrong field names, fenced |
Forced tool_choice | Ignored | Untestable — Ollama's API has no forcing mechanism |
Every behavior that's actually comparable matches, including the specific way the schema test fails. That's not proof by itself — lots of models get 4/4 on basic arithmetic — but it's one more independent line pointing the same direction as the tokenizer test, run through a completely different path (Ollama Cloud, not OpenRouter; the model's own name, not a stealth badge).
So: this is MiniMax. That part I'll put a number on — call it 97%, on the strength of a reproduced 36/36 primary-source match. Specifically M3.1-Flash-Preview is a weaker claim than that, and it deserves its own number, not the same one borrowed from the tokenizer test.
Here's why those two claims aren't equally certain. The tokenizer match proves the family, not the specific checkpoint — and now I have the data to show exactly how little it discriminates within that family. Six separate published MiniMax tokenizers, spanning M1 through M3, all tokenize these nine probes identically to space-bunny-alpha. That's strong evidence it's MiniMax. It is zero evidence for which MiniMax: the test can't tell M3.1-Flash-Preview apart from M1, M2, or anything else sharing that vocabulary, because none of them differ on the thing I'm measuring. It rules out Llama, GLM, Qwen, and DeepSeek completely; it rules out nothing inside MiniMax's own lineage. What points at M3.1-Flash-Preview specifically is a separate, weaker line of evidence: three spec matches (context size, mandatory reasoning, the same five-tier effort ladder) that aren't generic across this gateway, plus a reported reason the real checkpoint wouldn't be independently callable to test. That's a reasonable case, not a proof. I'd put it at 70–75% — good enough to run with, nowhere near the confidence behind "it's MiniMax."
Free is a launch price
The last two posts in this series had to argue that "free" was temporary by inference — a pricing row that read $0.00 with an expiration date decades out. This one doesn't need the argument. expiration_date: 2026-10-05 is sitting right in the registry response, three days from today. Whatever this is testing, the test has an end date, right in the API response.
That changes the advice slightly from the last two posts. It's not "architect so the free price isn't load-bearing." It's "assume this stops working on October 5, because the registry is telling you exactly when." The engine underneath — the 975k-token honest retrieval, the clean parallel tool calls, the fast cold start — may resurface under a different name, a different price, or not at all. Not every envelope problem is gateway plumbing, either: the schema failure reproduced identically on the real, named minimax-m3:cloud, so that one travels with the weights and won't get patched away by OpenRouter. The forced-tool_choice ignore and the disguised 502s are still plausibly route-specific — I couldn't test forced tool_choice against the named model at all, since Ollama's API doesn't expose that mechanism, and the 502-in-200 behavior is specific to this gateway's own error handling by definition. Don't bet on any of it in either direction. Bet on what you measured.
If you're running it today
- Never trust
response_formatwith this model, on any route. This one reproduces on the real, named MiniMax-M3, not just the stealth wrapper. Ask for JSON in the prompt by name, parse it yourself, validate the shape, reject loudly on a mismatch — the content is usually right, the field names won't be. - Treat
tool_choiceas a strong hint, not an enforcement mechanism. Check that the forced tool was actually called before trusting that it was. - Check for a nested
errorkey even on HTTP 200. This is the one most pipelines get wrong by default: status-code-only error handling will silently swallow a fake success with null content. Measured rate across three sessions: 3/48 (6.2%), clustered in one session, not steady. Check forchoicespresence, not just status code, regardless of how rare it looks on any given day. - Cost and capacity-plan off raw
completion_tokens, notcompletion_tokens_details.reasoning_tokens. The latter reports zero regardless of how much real, visible thinking happened. - The context window checks out, up to a point. Clean, honest retrieval through 974,950 reported tokens, and a real 400 past the 1,000,000 ceiling. Budget under a million. This is the first of the three stealth models in this series where the advertised ceiling and the measured ceiling are the same number.
- It's gone in three days. Don't build anything into this preview window that you can't unplug by October 5 without a scramble.
The whole thing in five lines
# the context claim is the most honest of the three stealth models so far
974,950 tokens clean, honest NONE on the phantom needle, then a real 400 at 1,000,000
# the envelope still drops two contracts
response_format -> silently reshaped; confirmed on named minimax-m3:cloud too, so it's the model, not the route
tool_choice -> ignored, not enforced
# the gateway lies about its own health, 3/48 across 3 sessions, bursty not steady
HTTP 200 + {"error": {"code":502, ...}} looks like success to anyone checking status codes only
# the lineage, with numbers instead of "almost certainly"
this is MiniMax: 97% -- reproduced 36/36 tokenizer match against MiniMax-M3's real weights
specifically M3.1-Flash-Preview: 70-75% -- 3 spec matches, not a tokenizer file for that checkpoint
first pass guessed Llama and was wrong; cellcog.ai's MiniMax call was right
# the clock is the headline
expires 2026-10-05 -- don't guess, inspect, and don't build what you can't unplug in 3 days
Related reading
- openrouter.ai/stealth/space-bunny-alpha — the model itself;
expiration_datestill reads 2026-10-05 as of this post - cellcog.ai: What Is Space Bunny Alpha? — the one outside writeup I found on this, pointing at MiniMax's M3.1-Flash-Preview; I verified both its tokenizer claim and its spec comparison independently rather than citing it on faith
- Union Alpha on OpenRouter: An Honest Scorecard for the Free Stealth Model — early drafts of this one wrongly inherited ox-alpha's identity guess before the real battery ran; this post is built to avoid that mistake a second time
- OpenRouter Handed Me a Free Mystery Model. I Put It Through a Real Scorecard. — the first of this series, on ox-alpha
- Running Claude Code Against Anything: Profiles, Other Models, and a Deliberately Raw One — the harness underneath all three of these probes
No comments:
Post a Comment