See more industry reports and analysis at signal65.com

Claude Fable 5.1 on Signal65 PINNACLE.

Nearly half the agentic errors, nearly twice the price of a correct answer.

One day after launch, Anthropic's new flagship posts a 7,898 on the Signal65 PINNACLE Model Score, 42% fewer weighted errors than Claude Opus 5, leads all five enterprise personas and completes every governed-data job in the suite. It also costs $2.46 per correct task, the most expensive correct answer from any model above the baseline. What the number means, and what it does not.

Signal65 · Ryan Shrout · Data basis: Claude Fable 5.1 scored September 2nd, 2026 over the Anthropic API in thinking mode; roster of 45 configurations through September 2nd, 2026


Signal65 PINNACLE launched on Monday with 44 model configurations, and the first question after any benchmark launches is what happens when the next frontier model lands. Claude Fable 5.1 shipped from Anthropic on September 1st, we ran it through the full suite on September 2nd, and the results are live on the results page today. It is the most capable configuration we have measured by a wide margin, it is the first to lead every one of the five personas, and it is priced accordingly. This post sets out what it did, what the score is and is not, what the remaining failures look like in a deployment, and where it lands on the cost of correct work.

The short version

Claude Fable 5.1 is the new top of the Signal65 PINNACLE table, with 42% fewer weighted errors than Claude Opus 5, the previous leader. That is a 1.73× error ratio, and about one-eighth the weighted errors of the Gemma-4-31B open baseline. It completed 276 of 280 agentic jobs under the strict whole-task rule, including all 160 governed-data jobs, and it invented an answer to 0.7% of questions the documents could not answer, against 7.6% for Opus 5 and 10.3% for GPT-5.6 Sol.

The score is a ratio of errors, not a grade out of ten thousand. 7,898 means the baseline model makes 7.9× the weighted errors Fable 5.1 does. It does not mean Fable is 7.9× smarter or does 7.9× the work, and because the scale is reciprocal, the last few points of accuracy at the top produce the largest jumps in the number. That is why the results views now default to a log axis, on which Fable sits one step above Opus 5 rather than towering over the chart.

The best model we have measured still fails about one workflow in ten on the hardest role. Per-step reliability compounds across a chained workflow, so a model that completes 98.6% of single jobs finishes about 90% of five-step IT Professional workflows once the retrieval steps and the compounding are counted. Fable 5.1 fails 7 of every 100 workflows across the five roles where Opus 5 and GPT-5.6 Sol fail 14, and a workload that runs ten thousand jobs a month at Fable's rate still hands someone downstream about 140 failures to catch.

It extends the cost frontier upward without bending it. At $2.46 per correct task on the full API bill, Fable 5.1 costs 1.8× Claude Opus 5 for 1.73× fewer weighted errors, so the top of the hosted market now charges for capability at almost exactly the rate it delivers it. A hundred dollars buys 41 correct tasks from Fable 5.1, 74 from Opus 5, 191 from GLM-5.2 on a leased B300 node and 1,295 from Qwen3.8-27B. The premium is the price of the last seven failures per hundred, and whether that is a bargain depends on what a failed workflow costs the business.

Five personas, one leader, and the same buyer's question five times. Fable 5.1 takes Knowledge Worker, Data Analyst and Executive from Opus 5, IT Professional from GPT-5.6 Sol and Customer Operations from GLM-5.2, with its widest margin on IT Professional, where chained tool use decides the score. Its advantage is concentrated in tool use, and on the retrieval-weighted roles the margin is narrower and the price gap is not, so GLM-5.2 remains the correct-work-per-dollar pick for a service desk on organized data.

Fig. 01: Claude Fable 5.1 on Signal65 PINNACLE, at a glance.

Claude Fable 5.1 on Signal65 PINNACLE, at a glance. Scored September 2nd, 2026.

Read the method first

Signal65 PINNACLE: Measuring Correct Work. A Methodology for Benchmarking Agentic AI in the Enterprise. Chapters 8 and 9 for the two data environments, 10 for fabrication, 11 for the personas, 17 for how to read each lens and 19 for what the benchmark does not claim. Every number below rests on choices that document explains.

The sections below run in the order a buyer should read them.

  1. What Claude Fable 5.1 did, and where the record comes from.
  2. What 7,898 means, and what it does not.
  3. The reality of agentic error rates at the top of the table.
  4. The price of the last step on the cost frontier.
  5. Five personas, one leader.

1. What Claude Fable 5.1 did, and where the record comes from

Every model on the Signal65 PINNACLE table runs the same jobs. The agentic half is 280 samples of enterprise tasks that run for 80+ tool-use rounds inside an environment generated fresh for the run, with the answer key generated alongside it and held outside the agent sandbox, graded in code with no model judging the output. 160 of those samples run against data an enterprise has already organized and governed, and 120 against data as it accumulates, duplicated, half migrated and contradicting itself. The retrieval half, RIKER, measures how well a model finds the right document at 128K context, how well it combines across documents, and how often it invents an answer the documents do not contain. Claude Fable 5.1 went through the whole suite in thinking mode over the Anthropic API on September 2nd, in our testing, one day after the model became available.

Fig. 02: Roster Widget Log

The live results view, ranked by PINNACLE Model Score with thinking builds only and the log axis that is now the default. Explore it on the results page.

The headline is the score, 7,898 on the PINNACLE Model Score, which reads as 42% fewer weighted errors than Claude Opus 5, half the weighted errors of GPT-5.6 Sol and 59% fewer than GLM-5.2, the best open-weight configuration. Underneath the score, Fable 5.1 completed 276 of 280 agentic jobs under the strict whole-task rule, and the four it missed were all on the as-found data. On governed data it completed 160 of 160, the first clean sheet on that half of the suite by any model in the roster, and on as-found data it completed 96.7%. Its fabrication rate of 0.7% is the lowest of any hosted model and the second lowest on the table behind Qwen3.5-397B-A17B with reasoning on, and its RIKER retrieval score of 97.3 is the best in the roster, ahead of Qwen3.8-27B, which held that title at launch.

Fig. 04: Where the record comes from

Where the record comes from. Claude Opus 5 still leads on as-found data and on aggregation; the new score rests on the governed clean sheet and on fabrication.

The composition of the record matters as much as its size, because the record does not come from the hardest condition. On the as-found data, the half of the suite that looks like an enterprise estate before anyone has cleaned it, Claude Opus 5 still completes more jobs than Fable 5.1, 98.3% against 96.7%, and Opus 5 also still leads on aggregation across documents, 97.0 against 95.3. Fable 5.1's 1.73× error ratio over Opus 5 is built on two things Opus 5 does badly. Opus 5 completes only 92.5% of governed-data jobs, the unusual case of a model that scores higher on messy data than on clean, and Fable 5.1 closes that gap entirely, and Opus 5 fabricates on 7.6% of unanswerable questions where Fable 5.1 holds to 0.7%. Both of those are exactly the failures the persona weights punish hardest, since Customer Operations weights fabrication at 40% and every role weights whole-job completion, so the score moved a great deal on a change that a single-condition accuracy table would have reported as a point and a half.

Fable 5.1 is no more token-efficient than its predecessor, at 100.7 correct answers per million output tokens against 103.4 for Opus 5, so it writes about 9,930 output tokens per correct job, hidden reasoning included, roughly the same as Opus 5 and more than twice GPT-5.6 Sol at 4,359. And this is one run of one hosted endpoint, made 24 hours after the model shipped, under the same resolution limits as every other row on the table. Hosted models can change behind an API without a version number changing, we will re-run Fable 5.1 as the roster refreshes, and a difference of a few points between two hosted frontier models on a single measurement is inside the noise the whitepaper describes.

Fig. 10: Model score vs. token efficiency

The live results view's scatter, Model Score against token efficiency on log axes. Fable 5.1 sits directly above Opus 5 at the same output-token efficiency, and the three GPT-5.6 tiers still sit alone on the right. Explore it on the results page.

42% fewer weighted errors than Claude Opus 5, and the record rests on governed data and fabrication, not on the messy half.

2. What 7,898 means, and what it does not

The PINNACLE Model Score is a ratio of errors. For each persona the benchmark adds up a model's error rates on the five things it measures, weighted by what that role fails on, and divides the baseline model's weighted error by the model's own, scaled so that Gemma-4-31B with reasoning on sits at 1,000 by definition and every other model is placed by how many fewer weighted errors it makes. The overall score is the geometric mean across the five personas, which punishes a model that aces one role and fails another. The persona weights are an editorial statement of what each role fails on, and they are the part of the benchmark most likely to be revised in future versions as the workloads and the personas grow, with any change versioned and published under the governance framework. Read that way, 7,898 says the open 31B baseline makes 7.9× the weighted errors Fable 5.1 does, or that Fable's weighted error budget is 12.7% of the baseline's, and 4,558 says Opus 5's is 21.9%. The gap between the two, nine points of weighted error, is a 1.73× ratio, and the note under the ranked chart on the results page states it the same way.

Fig. 03: Error ratio vs. score

The same six configurations on the two rulers. Nine points of weighted error between Opus 5 and Fable 5.1 is 3,340 points of score.

What the score does not say is as important on a day when the number at the top of the table nearly doubles. Fable 5.1 is not 7.9× smarter than a 31B open model, it does not do 7.9× the work, and it does not complete 7.9× as many jobs, since the baseline itself completes more than half of them. The reciprocal scale has a property that every reader of the table should carry, which is that each halving of the remaining error doubles the score whether the halving happens at 40% error or at 2%. That is deliberate, because at the top of the table the last error is the one that decides whether a chained workflow finishes, and a benchmark that flattened the top would hide the difference that matters most in a deployment. It also means the linear chart stretches at the top in a way the accuracy table does not, and a model whose weighted error was half of Fable 5.1's would score about 15,800 on the same ruler.

We spent part of this week deciding what to do about that, and the answer is in the results views today. The published number stands as it is, because it is an honest measurement, and what changed is the picture. The Model Score axis in both results views now defaults to a log scale, with a Log/Linear toggle on every axis so a reader who prefers the linear view is one click from it, and on the log axis equal distances are equal multiples of fewer errors. Fable 5.1 sits one step above Opus 5 on that axis, a step of 1.73×, which is a larger step than Opus 5 took over GLM-5.2 at 1.40× and a much smaller one than GLM-5.2 took over the baseline at 3.26×. The guidance for anyone quoting these numbers is the one we follow ourselves, which is to state a comparison as an error ratio or as a percentage of fewer weighted errors, and never as a model being some percentage better.

7,898 is a ratio of errors, not a grade. The baseline makes 7.9× the weighted errors. That is the whole meaning.

3. The reality of agentic error rates at the top of the table

A score at the top of the table invites the conclusion that agentic work is solved, and the arithmetic of the remaining failures says otherwise. Fable 5.1 missed four jobs in 280, all on as-found data. A single-task pass rate of 98.6% is the best we have measured, and the unit an enterprise deploys is not a single task but a chain of them, with every step needing to succeed for the workflow to deliver. Signal65 PINNACLE carries that compounding into a workflow-completion figure for each persona, measured per-step reliability raised to the number of steps a real workflow for that role chains, two for Knowledge Worker and Executive, three for Customer Operations, four for Data Analyst and five for IT Professional, and the results page publishes it beside every score. Those step counts are stated assumptions rather than measurements, and the whitepaper works the arithmetic in full, but the shape of the result does not depend on the exact integers.

Fig. 05: Half the failures, and still one in ten

Failed workflows per hundred started. Claude Fable 5.1 halves the failures of the previous leaders, and one in ten five-step IT workflows still does not complete.

On that basis Fable 5.1 completes 96% of two-step Knowledge Worker workflows, 94% of three-step Customer Operations workflows, 90% of four-step Data Analyst workflows and 90% of five-step IT Professional workflows, in our testing, for an average of about 7 failed workflows in every 100 across the five roles. Claude Opus 5 and GPT-5.6 Sol each fail about 14 in 100 on the same basis, GLM-5.2 about 16, and the Gemma-4-31B baseline fails 83 of every 100 five-step IT workflows. Halving the failure rate of the previous best models in one release is the real content of the record, and it is a larger practical change than the score suggests in one direction and a smaller one than the score suggests in the other, because the failures that remain are still counted in whole workflows that someone has to notice.

A workload that runs ten thousand agentic jobs a month at Fable 5.1's completion rate produces about 140 failed jobs a month, and on a five-step workflow the count of failed chains is closer to a thousand. Fabrication at 0.7% means seven invented answers in every thousand questions the documents cannot answer, and an invented answer arrives in the expected format in the expected place and flows into the next step, which is why it is the failure enterprises ask about most. None of this argues against the model, since every number above is the best in the roster, and all of it argues for the same deployment shape the launch article described, agents inside a workflow with a check at the end of the chain, sized to a failure rate that is now measurably smaller but is not zero. The difference between a benchmark record and a production standard is still measured in incidents.

The best model we have measured still fails one in ten five-step IT workflows.

4. The price of the last step on the cost frontier

Hosted models carry their full API bill, input and cached input and output at published list rates, with the tokens spent on failed attempts billed to the tasks that finished, and open-weight models are priced as a rented 8-GPU B300 node at $61 per hour, the median on-demand rate across GPU cloud providers researched September 2nd, scaled by the share of the node-hour the work uses. Claude Fable 5.1 lists at $10 per million input tokens and $50 per million output, double Opus 5 on both, with cache reads at $0.25 per million, which is 2.5% of the input rate where every other hosted model on the table charges 10%. On this workload that works out to $2.46 per correct task, $1.97 of it input and $0.50 output, so input is 80% of the bill, as it is for every hosted model on agentic work where the agent re-reads its context on every round. The cache-read price matters most on the retrieval half, where nearly all of the input is cached, and it is the reason the bill is not higher still.

Fig. 06: Model Score against cost per correct task

The live cost view at 50% node utilization rather than a best-case 100%, with the Fable 5.1 detail panel open. The last step on the frontier is Opus 5 to Fable 5.1, 1.8× the price for 1.73× fewer weighted errors. Explore it on the results page.

Fable 5.1 lands on the Pareto frontier, since no model shown is both cheaper and more capable, and it extends the frontier upward by a 1.73× error ratio for 81% more money. The frontier's shape is what it does not change. Between GPT-5.6 Sol, Opus 5 and Fable 5.1 the price of a correct task and the error ratio now move almost in lockstep, 1.12× the price for 1.15× fewer weighted errors from Sol to Opus 5 and 1.81× for 1.73× from Opus 5 to Fable 5.1, so the top three hosted models charge for capability at very nearly the rate they deliver it, and the days when Sonnet 5 offered 96% of Sol's score at half the price look like the exception rather than the rule at this tier. Below the hosted tier the frontier still runs steeply the other way. From DeepSeek-V4-Flash on a leased node at 14 cents per correct task to Fable 5.1 is 17.5× the price for 2.75× fewer weighted errors, and from Qwen3.8-27B with reasoning on at about eight cents it is 32× the price for 3.2× fewer.

Fig. 07: What $100 of correct work buys

What $100 of correct work buys, across the leading hosted and open-weight configurations above the baseline. Node-priced models at full utilization.

Stated as what a hundred dollars buys, Fable 5.1 delivers 41 correct tasks, Opus 5 74, GPT-5.6 Sol 82, Claude Sonnet 5 161, GLM-5.2 on a leased B300 node 191, DeepSeek-V4-Flash 710 and Qwen3.8-27B 1,295, at full utilization for the node-priced models, and the leased-node figures fall by half at 50% utilization while the API figures do not move. Fable 5.1 is the most expensive correct answer on the table from any model that clears the baseline. The one configuration above it, Qwen3-235B-A22B-Instruct at about $6, is expensive because it fails most jobs and bills for the attempts, while Fable is expensive because of its list price, and the distinction matters because one of those can be fixed by choosing a better model and the other is the going rate for the best one. As a rule of thumb, the roughly $1.10 premium per correct task over Opus 5 buys about seven fewer failed workflows per hundred, and whether that is a bargain depends entirely on what a failed workflow costs the business, which for a service desk, a compliance review or a production change is usually far more than the premium. The reading order the results views were built for still holds, which is to pick the intelligence band the job needs first and then buy that band at the lowest price per correct answer, and for the top band there is now one entry.

42% fewer weighted errors. 81% more per correct task. The top of the market now charges for capability at the rate it delivers it.

5. Five personas, one leader

The persona layer weights the same five measurements by what each role fails on, and at launch it produced three different leaders, Claude Opus 5 for Knowledge Worker, Data Analyst and Executive, GPT-5.6 Sol for IT Professional and GLM-5.2 for Customer Operations. Fable 5.1 leads all five, the first model to do so, with error ratios over the previous leader of 1.79× for Knowledge Worker, 1.47× for Data Analyst, 1.59× for IT Professional, 2.00× for Customer Operations and 1.59× for Executive, in our testing. The IT Professional score of 12,701 is the first five-figure persona score on the table, and it is the persona that weights agentic tool use at 55%, which is where Fable 5.1's clean sheet on governed data counts most.

Fig. 09: Five jobs, one leader

Five jobs, one leader. Ratios are the previous leader's weighted error divided by Claude Fable 5.1's.

The spread across its own five scores says where the advantage lives. Fable 5.1's best persona is 2.4× its worst, a narrower profile than Opus 5 at 3.4× or GPT-5.6 Sol at 4.4×, and a wider one than the open generalists GLM-5.2 at 1.4× and Qwen3.8-27B at 1.3×. The pattern behind the spread is that its lead is concentrated in tool use and fabrication, since on grounding it sits at parity with Opus 5 and on aggregation it trails it, so the roles that weight agentic completion most, IT Professional and Data Analyst, see the largest scores, and the roles that weight retrieval most see the smallest. Customer Operations is Fable 5.1's lowest persona and it still leads the role with half the weighted errors of GLM-5.2, because a 0.7% fabrication rate against 1.5% is the difference the persona is built to reward.

Fig. 08: Customer Operations Agent score by build

The live results view on the Customer Operations persona. The persona picker sits above the ranked bars on the results page.

At launch we wrote that the best model for customer operations agents was GLM-5.2, not Claude and not GPT, and on capability that sentence is now out of date after two days, which is what a living benchmark is for. The buyer's version of the sentence survives. GLM-5.2 with reasoning on delivers a correct task for 52 cents on a leased B300 node against $2.46 from Fable 5.1, it fabricates on 1.5% of unanswerable questions, and its 18-point drop from governed to as-found data still makes it the pick for an estate that has been organized rather than one that has not. Fable 5.1 is the model for the service desk where an invented policy is the failure the business cannot afford and the volume is low enough to carry the bill, GLM-5.2 remains the correct-work-per-dollar pick for the same role at scale, and the persona table is published so that a buyer can make that call by row rather than by headline. The same logic runs through the other four roles, with the constrained buyer's answer unchanged, since Qwen3.8-27B with reasoning on is still the best sub-50B configuration on every persona at about eight cents per correct task.

The best model for customer operations is now Claude Fable 5.1. The best value for it is still GLM-5.2.

What we are watching

Fable 5.1 will be re-run as the roster refreshes, because a hosted model is a moving target and one run on launch day is a data point rather than a verdict, and any movement will publish with the date beside it under the governance framework's correction and refresh rules. The scale is the second thing we are watching. The log-axis default and the per-axis toggles are live today, the published number stands as it is, and we are weighing a complementary quick-look figure for general readers that states correct work delivered rather than an error ratio, which would sit beside the Model Score rather than replace it. The third is the rest of the roster, because the open-weight side of the table has not had its answer to this release yet, the newest open-weight models are queued, NVIDIA's Nemotron family among them, and the AMD Instinct MI355X results publish when that system is characterized.

Claude Fable 5.1 is the best agentic model Signal65 PINNACLE has measured, by a margin that shows up in whole workflows rather than points, and it is the most expensive correct answer on the table from any model worth buying. Both of those are true at once, the benchmark exists to publish them side by side, and the results page carries every number in this post with the toggles to read it either way.


The results views, the methodology whitepaper and the governance framework are live at pinnacle.signal65.com. The launch article, Correct work, measured., covers the first data set in full. Claude Fable 5.1 results reflect one run on September 2nd, 2026 over the Anthropic API in thinking mode; API list rates researched September 2nd, 2026; B300 node rate is the median on-demand price across ten providers on September 2nd, 2026.