GPT-5.6 Sol, Terra, Luna: Pick the Tier, Then the Effort

GPT-5.6 Sol, Terra, Luna: Pick the Tier, Then the Effort

GPT-5.6 comes in three tiers, Sol, Terra, and Luna, and each has a reasoning-effort setting from low to max. That is twelve configurations. Artificial Analysis's data shows the choice is simpler than it looks, as long as you compare the right cost.

On price per token, a cheaper tier at max effort beats a pricier tier at low effort almost every time. On cost per task, it often doesn't, because max effort writes several times more tokens. The practical rule is to choose the tier by the quality you need at max effort, and only lower effort once you have measured that quality holds on your task.

The data

Artificial Analysis Intelligence Index v4.1 scores as of early July, with launch list prices per million input/output tokens and output speed:

ConfigIndexPriceSpeed (tok/s)
Sol · max59$5 / $3078
Sol · high56$5 / $30n/a
Terra · max55$2.50 / $15144
Luna · max51$1 / $6204
Sol · low49$5 / $3067
Terra · high49$2.50 / $15122
Terra · medium46$2.50 / $15139
Luna · high46$1 / $6237
Terra · low40$2.50 / $15135
Luna · low33$1 / $6227

Artificial Analysis has since moved to Index v4.3, which rescales every score, so compare these numbers with each other, not with current AA pages.

View 1: price per token

Rank by index score and per-token price and a clear pattern appears:

  • Luna · max (51, $1/$6) outscores Sol · low and Terra · high (both 49) at a fifth and two-fifths of their prices.
  • Terra · max (55) outscores Sol · low (49) at half the price.
  • Sol · high (56) is only one point above Terra · max (55) at twice the price.

By this measure, dropping effort within a tier is never the best move; dropping a tier and keeping max effort is.

View 2: cost per task

Per-token price is not what you pay. Max effort is expensive because it writes more: across the index suite, Sol · max generated about 70M tokens, while Sol · low generated 6.6M, more than ten times fewer. At the same per-token price, Sol · low costs roughly a tenth as much per task as Sol · max.

So Sol · low against Luna · max is not a clear win for Luna. Luna's per-token price is five times lower, but if Luna · max writes several times more tokens than Sol · low, the per-task costs converge. Measured on intelligence, per-task cost, and speed together, none of the ten configurations dominates all the others. Each is the best choice somewhere.

There is also a quality caveat. In its predeployment evaluation, METR detected Sol cheating (exploiting evaluation bugs or disallowed strategies) at a higher rate than any public model it had evaluated. Treat high index scores as an upper bound and check outputs on your own graders.

The rule

  • Step 1: pick the tier at max effort. Start with Luna · max. Move up to Terra · max, then Sol · max, only when Luna fails your evals. Each step costs 2.5x to 5x per token for about 4 index points.
  • Step 2: lower effort only with evidence. For high-volume or easy tasks, try high or low effort on the tier you picked and compare quality and cost per completed task. Keep the lower setting only if quality holds.
  • Don't switch to a pricier tier and run it at low effort. Sol · low loses to Terra · max on quality and is rarely the cheapest way to get a 49-level answer.
Artificial Analysis GPT-5.6 Sol (max) intelligence, speed, and price card
GPT-5.6 Sol (max) summary card from Artificial Analysis, July 2026.

Takeaway

Pick the tier at max effort; lower the effort only when you have measured that quality holds. Per-token comparisons flatter max effort, and per-task costs are what show up on the bill.

Update: OpenAI has cut prices since launch, starting with Sol to $4/$20 on August 21. The ordering above is unchanged, but check current rates on each model page.

Related: Best Reasoning Models · Best Cheap Models

Sources