Back to feed

OpenDesign quality-cost ranking: 17 models, no single winner

The OpenDesign design arena scored 17 models: GPT-6.1 Sol leads at 90.1, while GPT-6 Luna delivers 83.1 points at a $0.016 cost. I opened the table row by row and worked out what each score really costs.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — e9c02deed78
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

I laid the 17-row table in front of me: GPT-6.1 Sol 90.1 on top, GPT-6 Sol 89.9 right behind, Claude Sonnet 5.5 third at 86.8. The test is a design arena: each model ships a finished showcase piece and judges grade the result. Cost sits in the same table: $0.382 and $0.297 for the Sol siblings, $0.574 for Sonnet. These first three rows all come from open-design.ai records, where every model lists its score next to its estimated per-artifact cost.

The story of the leading pair stretches into late September: GPT-6.1 Sol replaced GPT-6 Sol on the September 29 DevDay stage, listed at $2 input and $10 output, one fifth of flagship Astra pricing. The venturebeat.com write-up draws the same line: Astra-class agentic coding and computer use, with an Ultrafast tier reaching 300 tokens/s. The table feels honest here: a 0.2-point gap in this AI showcase costs 29% extra. On the benchmark side the Sol family measures around 48 on the Intelligence Index, running near 100 tokens/s.

The Sonnet 5.5 camp arrived September 28: $2 input, $10 output, cache reads halved to $0.10. The artificialanalysis.ai measurement puts Sonnet 5.5 at 56 on the Intelligence Index, two points behind flagship Opus 5.5, while noting 70.6% on Terminal Bench 4.0 ahead of Opus. For this model , scored 86.8 in the design arena, the artifact bill reads $0.574. My read: for a team staying inside one house, Sonnet is the safe harbor and the bill stays predictable. The cache cut trims agent workloads by roughly 20%.

The top is pricey, the middle is smart

The middle rows are the most fun part of this table: MiMo 2.6 Flash at 83.5 for $0.029, GPT-6 Luna at 83.1 for $0.016, Grok 4.7 at 82.9 for $0.662. The openrouter.ai comparison page lays the structural gap bare: Grok 4.7 lists $2 input and $6 output over 500K context, while MiMo Flash lists $0.10 input and $0.28 output over 1.05M context. So a similar benchmark band with more than a 20-fold price gap. For showcase trials I would try the open-weight MiMo side or Luna, not Grok. Token spend decides here.

The arrowed row in the image is Claude Haiku 5.5: 82.2 points, $0.018. The anthropic.com announcement is dated October 7: $0.10 input and $0.50 output up to 100K tokens, both lines quintupling above it, 1M context. The artificialanalysis.ai reading places Haiku 5.5 second in its price class at max effort with Intelligence Index 43, costing $0.21 per task. The beam.ai review points the same way: 34 points at medium effort for $0.05, a made-to-measure role for sub- agent work. To me this row is the star of the table: 4.6 points under Sonnet at one thirtieth of the bill. That ratio is rare in LLM selection.

The Astra row breaks the pattern: 82.7 points, $1.61. The flagship model , listed at $10 input and $50 output, sits 0.4 behind Luna here at 100 times the bill. The Fable 5.1 row is even more striking: 80.3 points at $3.66, the priciest cell in the table. The artificialanalysis.ai records show Fable at $10 input and $50 output costing $2.37 to $3.76 per Intelligence Index task. My verdict is blunt: general AI power may be high, but in this arena it lands far off the price-performance line. I look at finished-job bills, not list prices, when I pick a model .

The expensive wing and the lower rows

The Flash corridor continues: DeepSeek V4.1 Flash at 81.2 for $0.023, GLM-5.3 Flash-X at 80.1 for $0.098. The yottalabs.ai comparison gives the backstage view: DeepSeek at 552B parameters in a $0.044 input and $0.30 output band, GLM in a $0.036 input and $0.50 output band. Live openrouter.ai prices point the same way. These two rows are my backup for the Haiku-Luna pair: if one stalls, the other steps in. The benchmark gap is 1.1 points, the bill gap 4-fold; here too the smart pick beats the pricey one. Cache-read cost rules the agent loop.

One step down sit GPT-5.6 Sol at 77.6 for $0.544, Hunyuan H4 Preview at 74.6 for $0.418, Qwen 3.8-Max at 72.0 for $0.595. The open-design.ai page hands its best-value badge to Luna: 92% of the top score at 4% of the top cost. That ratio carries the thesis I built all along: a model 5 to 13 points behind costs 30 to 40 times less. Qwen at a $2 input and $6 output list stays pricey in this arena. Paying per token and paying per finished job are different things; the table teaches that.

The final three rows close the table: Gemini 3.8 Flash 68.9 at $0.210, Muse Spark 1.3 66.6 at $0.267, Kimi K3 65.7 at $0.614. The gap to the top runs 21 to 24 points, with bills 13 to 38 times Luna. I read these rows as a warning, not a rejection: names that shine on general LLM leaderboards can look pale in a design showcase. When the task type changes, the benchmark order changes too. Not every model fits every showcase.

My selection logic

My rule condenses to three lines: I try Luna or Haiku 5.5 on short verifiable jobs, I stay with Luna past 100K tokens, and I bring Sol or Sonnet to the showcase final. The routing pattern in the beam.ai review says the same: the cheap model tries, a judge verifies, stuck work escalates to the big one. What the open-design.ai table taught me: customers rarely see the gap between 90 and 83 points, but they see the 24-fold bill gap every month. Skill in AI spend is picking the right tier, not the top score.

Score per cost: value belt vs top tier
MetricResult
Top scoreGPT-6.1 Sol 90.1
Value leaderLuna 83.1 / $0.016
SpotlightHaiku 5.5 82.2 / $0.018

Key moments

  1. Opening the 17-row table
  2. Top trio and the bill gap
  3. The value belt in the middle
  4. Haiku 5.5 spotlight and the 100K line
  5. The pricey wing: Astra and Fable
  6. Flash corridor and lower rows
  7. My three-line selection rule

AI commentary

“I stared at this chart for an hour and my view is firm: the gap between the top and the middle is closed with judgment, not money. Never send the expensive model at every job; the cheap tier now ships showcase work, and I pick accordingly.”

AI assessment

The strongest counter-view: this table measures one showcase, not general AI strength. A model scoring 90 in the design arena may not hold that rank in long-horizon agent coding or knowledge exams. On the artificialanalysis.ai Intelligence Index Sonnet 5.5 sits at 56 near the top while holding 86.8 and third place here; the two tests weigh different virtues. Treating this table as a general benchmark would be a mistake, and I avoid it.

The gaps are plain too: speed, latency, language coverage and long-context behavior are absent. Per-artifact cost is an estimate; my real bill moves with request counts, cache hits and output length. A Haiku request crossing the 100K line costs five times more, which no table cell shows. That is why the two-tier list on the anthropic.com page matters. I study footnotes and price pages, not just cells.

I note the possible interest as well: the open-design.ai page opens onto a download page serving 21 agents , with the best-value badge on Luna. That framing does not taint the measurement, but I never let one table decide. I cross-read venturebeat.com, artificialanalysis.ai and beam.ai records. My principle: one table sparks curiosity, three tables decide.

The takeaway for readers is simple: draft showcases with Luna, run volume production with Haiku 5.5 under a judge pattern, and bring Sol or Sonnet to the final jury. Watch the 100K token line and keep cache-read ratios high. Live openrouter.ai prices shift weekly, so check the list monthly. In this setup I protect both showcase quality and the month-end bill.

Sources

8 links; 2 of them also cited by 7 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

artificial intelligence · model · benchmark · agent · llm

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review source passages, versions and origins.

READ WITH SOURCES

Understand this story.

Checking your account…