🔍 Read the full analysis: The Roles Behind My September 2026 AI Stack on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
In a September 29, 2026 report, Thorsten Meyer describes using Claude Opus 5.5 for building and GPT-6.1 Sol for detailed work and review. His model choices draw on Artificial Analysis Intelligence Index v4.3.x scores and estimated task costs; he cautions that the index does not establish which model will perform best on a reader’s workload.
Thorsten Meyer said on September 29, 2026, that he uses Claude Opus 5.5 as his main model for building and GPT-6.1 Sol for detailed work and review. His account frames the choices around a reported gap between benchmark scores and estimated task costs: several models score within about 20 points on Artificial Analysis’s index, while their listed cost per task varies by roughly 100 times.
Meyer’s model comparisons use the Artificial Analysis Intelligence Index v4.3.x, which he describes as a general capability measure rather than a verdict on a specific workload. In his table, Opus 5.5 scores 58 at its top setting and costs an estimated $5.98 per task. GPT-6.1 Sol at xhigh scores 51 and costs $0.39 per task. The table lists Luna at $0.07 per task and a score of 37. These are index-based estimates, not a guarantee of what a particular user will pay or get.
Meyer assigns Opus 5.5 to features, APIs, multi-file work and refactoring, generally at high effort. He reserves xhigh for demanding work such as architecture, migrations and trust boundaries. He says the high setting scores 54 for $1.82 per task, while xhigh scores 56 for $3.46. The maximum setting scores 58 at $5.98; Meyer argues that its extra cost is rarely justified for his work.
For GPT-6.1 Sol, Meyer uses high or xhigh for focused investigation and review. He reports costs of $0.32 and $0.39 per task at those settings, respectively. He says he uses other models selectively: Astra or Fable for a second opinion if models disagree, Sonnet for scoped subtasks, and Luna for routine checks and bulk classification. He also describes Jev as a decision model for high-volume yes-or-no judgments and routing, but gives no benchmark or cost figures for it in the supplied account.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
How Meyer Splits Building and Review
The report’s practical point is that model selection can be a task allocation decision, as well as a ranking. Meyer’s workflow assigns the costlier, higher-scoring Opus to work that builds or changes software, while using Sol for repeated review and investigation. Based on his cited estimate, a Sol review pass costs a fraction of an Opus task, which could make routine second-model checks more affordable for his workflow.
Meyer argues that review by a different model family can catch problems a model may miss in its own output. He also sets limits on that argument: a second model may share a flawed specification, and passing tests alone does not mean work is ready to ship. The report presents these as operating rules and judgments, not as results from a controlled comparison of review accuracy.
For readers choosing models, the figures are a reason to measure their own tasks, not a universal ranking. Meyer advises shadow-testing before switching. The benchmark scores and estimated costs can help narrow candidates, but they do not establish how much human review a task needs or whether a lower model bill reduces total project cost.
Scores, Effort and Estimated Costs
Meyer’s September account compares six models: Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. He reports Opus ahead on the index, while Sol and Luna have the lowest listed task costs. The comparison places Sonnet at its maximum setting at $7.60 per task for a score of 56, versus Opus at maximum for $5.98 and a score of 58. Meyer uses that comparison to argue against running Sonnet at its most expensive setting in his own stack.
His effort-level table shows that the estimated cost changes with the setting. For Opus, the listed cost rises from $1.34 at medium to $5.98 at maximum, while the score moves from 51 to 58. For Sonnet, it rises from $0.59 to $7.60, with the score moving from 41 to 56. Meyer says medium remains his default for documents and everyday work, while Sonnet’s best value in his assessment is high: a score of 47 for $1.08 per task.
The report also gives token prices per million tokens: Opus at $4 input and $20 output, with cache reads at $0.20; Fable and Astra at $10 and $50; Sol at $2 and $10; and Luna at $0.10 and $0.50. These token prices and the reported per-task estimates describe different measures. A task’s token use affects its cost, so the per-task figures should not be treated as a fixed price for every request.
““the question” changes from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?””
— Thorsten Meyer
What the Index Cannot Settle
The source does not provide the underlying task-level benchmark methodology, the full set of prompts, or independent results from Meyer’s own work. It also does not establish that the estimated costs will match other users’ workloads. Meyer says one index point is within the noise, and notes that low and maximum settings for GPT-6.1 Sol had not yet been published at the time of his report.
His account reports that Sol’s high and xhigh settings take 57 and 69 seconds to produce a first token, respectively, which he says makes them unsuitable for interactive use at those settings. The figures are tied to the index tests; performance in a different product or task may vary. The report does not give measured review accuracy, failure rates, or total human time across the model assignments.
The supplied source ends partway through an illustrative comparison about human review erasing model-cost savings. It does not provide the example’s full numbers, so the size of that effect cannot be assessed from the material provided. The report also does not specify how Jev was evaluated or priced.
Testing Models on Real Work
Meyer’s stated next step for anyone considering a switch is to shadow-test candidate models against existing work before changing the default. That means comparing outputs against the quality bar for the task and accounting for the effort setting, estimated model cost and any additional human review. His report does not announce a follow-up test or give a date for one.
Further benchmark results may change the comparison: Meyer says some Sol effort settings were not yet listed, and the index scores and cost estimates are snapshots from September 29. For now, his published stack is a description of his own operating choices. Whether the same division works for other teams remains a question for their own task-level evaluations.
Key Questions
Which model does Meyer use for building?
He identifies Claude Opus 5.5 at high or xhigh effort as his main model for building, with xhigh reserved for harder problems such as architecture and migrations.
What role does GPT-6.1 Sol play in his stack?
Meyer uses Sol at high or xhigh for detailed investigation and review. He reports estimated costs of $0.32 to $0.39 per task for those settings, based on the index figures he cites.
Does the benchmark show which model is best for every user?
No. Meyer calls the Artificial Analysis Intelligence Index a map of general capability, not a verdict on a specific workload. He recommends testing models against the work they would handle.
Why does Meyer use different models for building and review?
He says a review by a different model family can provide another perspective on Opus’s output, and that Sol’s estimated per-task cost makes routine review affordable in his workflow. The report does not provide measured evidence that this improves review accuracy.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
