Only one outlet has reported this event so far — you're reading it below, credited to its source. AIPROPX is tracking the web for more coverage; as additional outlets confirm it, this becomes a full multi-source story automatically.
By AIPROPX Editorial Desk · Published · Updated
One outlet is reporting this so far. AIPROPX is tracking it and will gather every additional source as it develops — the full multi-source comparison appears automatically once a second outlet confirms it.

A 150M model reached 29.5% while costing just $0.0007 per task
ChatGPT scored higher, yet its comparable reasoning runs cost substantially more
BDH-CQ performs reasoning internally instead of generating lengthy intermediate text
Pathway, an AI lab focused on building Post-Transformer architectures, has released new benchmark results for its BDH-CQ reasoning model.
According to the researchers , their 150M-parameter model scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set.
It achieved this at a computed inference cost of $0.0007 per task, roughly eleven times cheaper than ChatGPT's underlying GPT 5.6 Luna (Low) model.
A cheaper way to reason
Today, many AI tools waste computing power because of how they are designed, not because deep reasoning demands it.
"Today's AI pays a steep token cost for reasoning, but that cost is imposed by architecture, not by any law of intelligence," said Zuzanna Stamirowska, CEO and co-founder of Pathway.
“We show that a different architecture changes the game and opens up a whole new space in terms of how much intelligence per dollar.”
Amazon Web Services believes that BDH-CQ’s result is a promising step toward using advanced AI reasoning in real products more affordably.
"Customers are increasingly exploring how to move advanced reasoning from experimentation into production, where performance, efficiency, and scalability all matter," said Nicolas Tarducci of AWS.
ARC-AGI-1, a widely used reasoning benchmark for AI systems, checks whether a system can infer an underlying rule from limited examples and apply it correctly to new inputs.
In this test, OpenAI's Luna model scored only slightly higher at 34.2%, yet running it still costs significantly more ($0.008 per task).
That price gap already includes OpenAI's recent 80% price cut on Luna, which began on July 30th of this year.
Further up the chart, Claude Opus 5 and Gemini 3.1 Pro reach 97–98% but cost around $0.5 – $0.6 per task, meaning the frontier's very top costs close to a thousand times more than BDH-CQ for the highest scores.
On the cheap end, Qwen3 235B costs over three times more than BDH-CQ while scoring worse than even its Low variant, so it isn't a real competitor on either price or performance.
"Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning," said Łukasz Kaiser, co-author of the original 2017 Transformer paper.
Why it costs so much less
The efficiency gap stems mainly from a structural difference in how each system actually performs reasoning during inference computations.
Many reasoning AI systems generate intermediate text, adding one token after another before producing their final answers.
The longer that written reasoning becomes, the more it costs to run and the slower the AI responds to each request.
BDH-CQ works quite differently, quietly solving problems inside its own memory instead of writing everything down first as visible text.
Pathway also confirmed that early experiments already follow standard Transformer-like scaling laws across model sizes from 1B to 600B parameters.
The company also plans to extend this approach toward harder benchmarks, including mathematical reasoning, ARC-AGI-2, and eventually full ARC-AGI-3 evaluations.
If these efficiency gains hold across larger and more difficult tasks, cost rather than raw capability could increasingly separate rival reasoning systems.
Indexed and credited by AIPROPX. Originating outlet: TechRadar. Open at source →
An original, deterministic readout — composed only from the computed coverage facts on this page. No interpretation, no rating; figures only.
AIPROPX has consolidated 1 report from 1 outlet into a single canonical entry on “11X cheaper than ChatGPT: Tiny 150M model just proved AI doesn't need to "think out loud" to be smart.” Every covered outlet is based in Other.
The only timestamped report came from TechRadar (Aug 13, 2026, 01:05 UTC).
Comparing the wording across sources, the phrase recurring most across the coverage is “150m model”.
2 statements are carried by only one outlet within this set and are not echoed by the others.
Every figure above is a direct count of real published articles. AIPROPX indexes and compares the original reporting — it never rewrites, rates, or editorializes — and each publisher’s full article is always one click away.
Generated by AIPROPX from the source counts above. AIPROPX indexes and resolves coverage; the original publishers are credited and linked at origin in every report.
Coverage from 1 independent outlet across 1 region — each view opens on its own page.
AIPROPX — “11X cheaper than ChatGPT: Tiny 150M model just proved AI doesn't need to "think out loud" to be smart” · https://www.aipropx.com/story/0b98246fc849d047b9c4e365f61fdf20
Other events being covered across multiple sources right now.
Inside Paramount, employees say they have bigger concerns than a reported plan to ditch California
3 outletsFlightAware dropped its lawsuit against Kalshi after just one day
2 outletsIPhone 18 Pro set to become Apple’s new default model
2 outletsWhich Trump Officials Joined Him on His Secret Flight Out of Turkey?
2 outletsJapan and the U.S. just spent billions to try to save the yen. Why is it already losing ground?
2 outletsGoogle Gemini Just Hit One Billion Monthly Users in Record Speed