AI is New Zealand's next export opportunity
At the Aotearoa AI Summit, Autohive co-founder JD Trask argued New Zealand should stop treating AI mainly as a governance and risk topic and start …
Read article
Pricing and specs in this post are current as of 23 September 2026 (NZST). Model rates change often. Check the linked vendor pages before you budget.
GPT-6 Sol, GPT-6 Luna, and Grok 4.7 are now live in Autohive. The price gap between them is big. Luna runs at $0.10 input and $0.50 output per million tokens, while Sol and Grok cost roughly 20 times more on input.
The short verdict: send heavy reasoning to Sol or Grok, and push high-volume routine work to Luna. We made the same call when GPT-5.6 landed. See the earlier GPT-5.6 routing guide for that breakdown.
Pick any of the three in Agent Creator, assign them per agent inside a multi-agent workflow, run them on scheduled jobs, or test them straight in workspace chat. All three support vision, tool use, reasoning, and streaming in the Autohive catalog. The model IDs are gpt-6-sol, gpt-6-luna, and grok-4.7.
Sol is the workhorse of the GPT-6 line. It scores below the flagship Astra model on raw intelligence, but it costs a fraction of the price. OpenAI lists it at $2.00 input, $0.20 cached input, and $10.00 output per million tokens. Context runs to 1.05M tokens, with a 128K max output.
Sol reads text and images and writes text back. It supports function calling, web search, file search, and computer use, and you can dial reasoning effort anywhere from none to max. Push past 272K tokens and the long-context rate jumps to $4 input and $15 output, so keep an eye on payload size for big jobs.
Artificial Analysis reports an Intelligence Index of 48 and a Coding Agent Index of 57 at max effort. Output speed lands around 126 tokens per second, though that number moves with your reasoning settings.
Luna is the price story of this release. At $0.10 input, $0.01 cached input, and $0.50 output per million tokens, it’s one of the cheapest capable models you can run today. It shares Sol’s 1.05M context window, 128K max output, and the same modalities, tools, and reasoning settings.
Above the context threshold, the long-context rate is $0.20 input and $0.75 output. Artificial Analysis reports an Intelligence Index of 37 standard, or 41 at max effort, and a cost of about $0.07 per test task. Reported output speed sits around 157 tokens per second.
This is a price-and-scale release. Sol and Grok still lead on hard reasoning tasks. Luna’s job is to run cheap, high-volume work at scale.
Grok 4.7 from xAI lists at $2.00 input, $0.50 cached input, and $6.00 output per million tokens, with a 500K context window. It reads text and images and writes text back, with function calling, web search, X search, and code execution. Reasoning effort runs from low to xhigh.
A faster variant exists at $4/$12, but it’s restricted to Cursor and Grok Build, so you can’t select it inside Autohive. The US regional endpoint adds a 10% token premium on top.
Artificial Analysis reports an Intelligence Index of 46, a Coding Agent Index of 56, and CursorBench at 46.3%. xAI’s own number for DeepSWE is 71% at high effort, worth treating as a vendor claim rather than an independent score. Artificial Analysis also reports a 29% hallucination rate and about 81K output tokens per task on average in its test setup, 125% more than Grok 4.6 and 196% more than GPT-6 Astra in the same tests. Average task duration ran around 7.1 minutes. The lower per-token output rate looks cheap on paper, but heavy token use can erase the saving fast.
| Model | Input $/M | Cached input $/M | Output $/M | Context | Max output | AA Intelligence |
|---|---|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | 1.05M | 128K | 37 (41 max) |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | 1.05M | 128K | 48 |
| Grok 4.7 | $2.00 | $0.50 | $6.00 | 500K | n/a | 46 |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | 1.05M | 128K | 53 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | 1M | 128K | 58 |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | 1M | 128K | n/a |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 200K | 64K | n/a |
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | 1M | 128K | n/a |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 under 200K | 1.048M | n/a | about 30 |
Notes: Sol’s long-context rate above 272K is $4/$15. Luna’s long-context rate is $0.20/$0.75. Grok’s Fast variant at $4/$12 is limited to Cursor and Grok Build. Anthropic offers an Opus 5.5 Fast tier at $8/$40. Gemini pricing shown is the under-200K rate.
Direct head-to-head data is thin because these releases are new. Treat every number as a signal and test on your own workload before you hand a model a critical job.
Independent numbers from Artificial Analysis:
xAI’s own number:
Reported output speeds put Luna around 157 tokens per second and Sol around 126 tokens per second. These come from different runs with different reasoning settings, so don’t read them as a controlled race.
Sol is the strongest of the three on general intelligence and coding. Grok 4.7 is close on coding but carries a hallucination and token-cost risk. Luna trails on hard reasoning and wins on price. Opus 5.5 leads the wider field on intelligence score, at a higher price.
These are modeled estimates. They exclude Autohive platform fees, tool charges, retries, and provider-specific caching rules.
| Model | Cost |
|---|---|
| Luna | $0.10 |
| Haiku 4.5 | $1.00 |
| Sol / Grok / Sonnet 5 / Gemini | $2.00 |
| Opus 5.5 | $4.00 |
| Astra / Fable | $10.00 |
| Model | Cost |
|---|---|
| Luna | $0.50 |
| Haiku 4.5 | $5.00 |
| Grok 4.7 | $6.00 |
| Sol / Sonnet 5 | $10.00 |
| Gemini | $12.00 |
| Opus 5.5 | $20.00 |
| Astra / Fable | $50.00 |
Assume 120K cumulative input and 6K output, uncached.
| Model | Cost per run |
|---|---|
| Luna | $0.015 |
| Haiku 4.5 | $0.15 |
| Grok 4.7 | $0.276 |
| Sol | $0.30 |
| Sonnet 5 | $0.30 |
| Gemini | $0.312 |
| Opus 5.5 | $0.60 |
| Astra / Fable | $1.50 |
Assume 6B input and 300M output tokens.
| Model | Modeled monthly cost |
|---|---|
| Luna | $750 |
| Grok 4.7 | $13,800 |
| Sol / Sonnet 5 | $15,000 |
| Gemini | $15,600 |
| Opus 5.5 | $30,000 |
| Astra / Fable | $75,000 |
| Model | Modeled monthly cost |
|---|---|
| Luna | $372 |
| Grok 4.7 | $7,500 |
| Sol | $7,440 |
| Opus 5.5 | $14,040 |
Caching changes the math, but cache-write costs and expiry can shift the real total. If Grok’s heavy output-token pattern shows up in your workload too, it can cost more than the sticker price suggests.
Moving routine, high-volume work to Luna instead of Sol cuts this modeled monthly bill by about 95%. That’s the case for tiered routing, in one number.
Autohive lets you mix models inside one workflow. A cheap worker can clean, classify, and trim data, then hand a smaller payload to a stronger reasoning model. Swapping a model doesn’t force you to rewrite an agent’s integrations or custom API wrappers.
| Job | Best pick | Why |
|---|---|---|
| Lead planner or orchestrator | Sol | Strong reasoning, big context, wide tool support |
| Coding and terminal work | Sol or Grok 4.7 | Sol leads on coding index; Grok is close but watch token use |
| Long technical documents | Sol | 1.05M context and 128K output |
| Complex tool chains, high-stakes analysis | Sol | Best of the three when a mistake is expensive |
| Classification, extraction, triage | Luna | Cheap enough to run at scale |
| Scheduled jobs and routine summaries | Luna | Runs on a timer without draining budget |
| Worker nodes in multi-agent flows | Luna | Does the bulk work and passes clean data upstream |
| Structured JSON, high-volume ops | Luna | Lowest cost per task by a wide margin |
| Research with current web or X context | Grok 4.7 | Web and X search plus code execution |
| Critique and long-form output | Grok 4.7 | Lower output rate can help if token use stays controlled |
A practical pattern: run a scheduled Luna job overnight to pull and classify incoming data, then trigger a Sol agent only on the items that need judgment. You pay Luna prices for the volume and Sol prices for the hard 5%. See the Job Scheduler docs and the multi-agent setup guide.
The price is identical at $2/$10. Pick based on your own task tests, tool needs, and provider fit.
Opus scores higher on the intelligence index, 58 versus 48, but costs twice as much on input and output. Use Opus when that extra score is worth the extra spend.
Astra costs $10/$50 and scores 53. Sol scores 48 at one-fifth the price, which makes Sol the practical default for most work.
Luna runs about 10 times cheaper on input and output, and it has a bigger context window too. Test both on your own data before moving production work over.
Grok’s output is cheaper at $6 versus $10, and its coding index sits one point behind Sol. Sol is the safer default because Grok’s hallucination rate and heavy token use can raise risk and cost. Choose Grok when X search or its research profile matters.
Both suit research work. Gemini’s output rate under 200K is $12 against Grok’s $6, and its reported intelligence score is lower too. Test your exact workflow before picking one.
How much does GPT-6 Sol cost? $2.00 input, $0.20 cached input, and $10.00 output per million tokens. Above 272K tokens, the rate rises to $4 input and $15 output. Rates current as of 23 September 2026.
How much does GPT-6 Luna cost? $0.10 input, $0.01 cached input, and $0.50 output per million tokens. Long-context rate is $0.20/$0.75. Rates current as of 23 September 2026.
How much does Grok 4.7 cost? $2.00 input, $0.50 cached input, and $6.00 output per million tokens. The Fast variant at $4/$12 is limited to Cursor and Grok Build. The US endpoint adds 10%.
GPT-6 Sol vs Grok 4.7, which is better? Sol edges ahead on coding and general intelligence in the independent scores cited above. Grok may fit better when X search or its research profile matters. Test both on your workload.
GPT-6 Sol vs Claude Opus 5.5? Opus scores higher but costs twice as much. Sol is the value default. Use Opus for jobs where stronger reasoning pays back the extra cost.
GPT-6 Luna vs Claude Haiku 4.5? Luna is about 10 times cheaper on input and output and has a larger context window. For high-volume routine tasks, Luna usually wins on cost.
Which model should I use in Autohive? Use Luna for routine work at scale, Sol for reasoning and coding, and Grok 4.7 when its research profile or X context matters. Mix models in one workflow to control cost.
Open Agent Creator, pick gpt-6-sol, gpt-6-luna, or grok-4.7, and test it in workspace chat. Then use tiered routing so Luna handles the volume and Sol handles the hard calls.
At the Aotearoa AI Summit, Autohive co-founder JD Trask argued New Zealand should stop treating AI mainly as a governance and risk topic and start …
Read articleClaude Code and Claude Cowork went down for nearly three hours on 28 August 2026, the latest in a string of Claude outages this year. Here is what …
Read article