Correct work, measured.
What the first Signal65 PINNACLE results say about agentic AI in the enterprise.
Forty-four model configurations, three accelerator platforms, two serving engines and six hosted APIs, measured on the same enterprise jobs and graded in code. The first Signal65 PINNACLE data set says what agents finish, what they invent, what one node holds and what a correct answer costs, and the ordering does not match any general capability leaderboard.
Signal65 · Ryan Shrout · Data basis: results through 30 August 2026
Enterprises are deploying agents that run for 80+ tool-use rounds, read policy documents, apply governing rules, query databases and hand a set of deliverables to someone downstream who depends on them. The numbers available to evaluate those deployments each answer half the question, since a capability leaderboard says whether a model can answer a question and an infrastructure benchmark says how many tokens a platform can move, and neither says whether the work came out right on the data an enterprise holds, at the volume it needs, for a price it can defend.
Signal65 PINNACLE measures correct work. Every job runs inside an environment generated fresh for that run, the answer key is generated alongside the environment and held outside the agent sandbox, and the result is graded in code, so no model judges the output and no human rates it. Every job runs twice, once against data an enterprise has already organized and once against data as it accumulates, and every model carries a measured rate for how often it invents an answer the documents do not contain. Model quality, platform capacity and the cost of a correct task come out of the same runs, so the cost is measured rather than modeled.
One measurement set, three lenses. Model intelligence, silicon capacity and solution economics come from the same scored runs.
This article is about what the first data set says. The method is set out in full in the methodology whitepaper, Signal65 PINNACLE: Measuring Correct Work, and the whitepaper is the companion to keep open beside these results, because every number below rests on choices it explains and defends. Forty-four model configurations from 30 base models and 12 makers, three accelerator platforms, two serving engines and six hosted APIs were measured on the same enterprise jobs through 30 August 2026 in our testing, and the roster grows from here.
Read the method first
Signal65 PINNACLE: Measuring Correct Work. A Methodology for Benchmarking Agentic AI in the Enterprise. Chapters 8 and 9 for the two data environments and the gap between them, 10 for fabrication, 11 for the personas, 13 for the concurrency sweep and the gates, 15 for precision, 17 for how to read each lens, 19 for what the benchmark does not claim.
The first data set at a glance, through 30 August 2026.
The takeaways run in four collections, and the sections below walk through them in order.
The partners behind the silicon results
A benchmark like this cannot be built alone, and our first outreach went to the primary GPU vendors, NVIDIA and AMD. Both teams have been generous with hardware, engineering time and advice as the method took shape, and we are proud of the response and the partnership we have gotten from both. NVIDIA delivered the Blackwell Ultra B300 system that anchors the launch speed reference, run exactly as delivered and not modified by the Signal65 AI Lab. AMD engineering has worked with us on the Instinct MI300X results published today, provided the MI355X system the lab began bringing online last week, and received the same materials for review. Both companies were given the methodology whitepaper and the governance framework for a technical read before publication.
"Agentic AI benchmarks must measure more than raw performance — they must show how much correct work a platform delivers, at what speed, and at what cost. This is exactly what the industry needs more of, and NVIDIA Blackwell Ultra serving as the reference platform for Signal65 PINNACLE speed measurements is a step in that direction. We look forward to working with Signal65 as coverage expands across the ecosystem."
Ian Buck, SVP for Hyperscale and HPC, NVIDIA
“Agentic AI is multi-turn, with growing context, KV-cache reloads, and tool calls. Enterprises need to measure correct work on real hardware, not just raw token throughput. Signal65 PINNACLE is built to answer that question, and AMD looks forward to supporting its evolution with Instinct and ROCm.”
Ramine Roane, Corporate Vice President of AI Applications, AMD
Partnership comes with published rules, and they live in the governance framework. Publication is unconditional. Vendors get a factual-accuracy review window before platform results publish, with no approval rights, no vetoes and no right to delay, every company working with Signal65 PINNACLE has the same level of input, and decisions about which workloads are included rest with Signal65, with the reasoning published. The same model scores the same across NVIDIA and AMD platforms in our testing, inside run-to-run noise, so a quality score cannot be moved by choosing a box. Vendors supply hardware, engineering time and input. They do not supply the answer. That input loop is also just beginning, since the governance framework invites each vendor to supply the configuration it would recommend to a customer running this workload, with results run and disclosed as vendor supplied, and that exchange is where we expect the platform numbers to move fastest from here.
1. Agentic work is a different test
The first collection of takeaways is about the work rather than any one model. On organized, governed data the top of the roster is close to saturating, with nine of 44 configurations completing at least 95% of jobs in our testing. On the same jobs run against data as an enterprise keeps it, duplicated, half migrated and contradicting itself, two configurations clear 95%, Claude Opus 5 at 98.3% and GPT-5.6 Sol at 95.8%, and the median configuration gives up about 28 points of workflow completion between the two conditions. Clean data is close to solved for the scenarios in this first suite, and the enterprise is not. The suite grows from here, with more job types and deeper workflows to come, and those will put the clean-data side of the story back under test. Either way, the gap between the two scores is a per-model property that maps onto a procurement decision, since a wide gap is a scope problem a buyer can engineer around by cleaning first and a low score on both sides is a capability problem they cannot.
Saturation also deserves a second read before anyone relaxes, because the ceiling on this benchmark is not a grading curve. Every environment is generated fresh alongside its answer key, so a model that can truly do this work should score 100% on every draw, and the particular entity names, file layouts and data values an environment happens to contain should change nothing. In practice they do. The best governed score in the roster is 99.4%, not 100, the model that leads the overall table completes 92.5% of governed jobs, and the misses move with the draw, in our testing. That residue matters at enterprise scale, since a workload that runs ten thousand times a month at 98% per run is two hundred failures a month for someone downstream to catch, so close to solved on a benchmark and solved in production are different standards, and the distance between them is measured in incidents rather than points.
The same job on both sides, with the constructions that punish a plausible shortcut. Sandboxes are generated per run; file names and values here show the shape of each trap.
Every configuration, both conditions. One line rises.
Claude Opus 5 is the line that does not fall, at 92.5% on governed data and 98.3% on as-found data, and the other 43 configurations drop by a median of about 28 points. The widest cliffs belong to models that look strong on clean data, with Qwen3.5-397B-A17B falling 64 points and Mistral-Medium-3.5 falling 51. A single-condition benchmark would have published the left column and stopped there.
Clean data is close to solved. The enterprise is not.
The second finding is arithmetic, and it explains why the first one matters more than a five-point difference suggests. Per-step success compounds across a chained workflow, so a model that is 90% reliable per step finishes a five-step job 59% of the time, at 95% it finishes 77%, at 98% it finishes 90% and at 99% it finishes 95%, and every halving of the per-step error doubles the length of workflow a model can carry at even odds. On the five-step IT Professional workflow GPT-5.6 Sol finishes 83% of the time in our testing, the Gemma-4-31B-it reference model about 17%, and the smallest model in the roster fewer than one time in a thousand. The compounding term is why the last point of accuracy at the top of the table is worth as much as five points in the middle of it, and it is built into the Signal65 PINNACLE productivity measures because a benchmark that grades one task at a time is grading the wrong unit.
Per-step reliability compounds across workflow depth. The whitepaper works the N = 5 example in full.
Tokens follow the same logic. The same agentic job costs GPT-5.6 Luna about 4,000 output tokens and Qwen3.5-9B about 46,000, an 11× range with hidden reasoning tokens counted for every model, and once the tokens spent on failed attempts are charged to the jobs that finished, the range runs past 100×, with the small fast models that look cheapest per token becoming the most expensive per correct answer. A tokens-per-second figure measures how fast a box emits work, not how much work a correct answer takes, and until the token count is attached to a graded job the price of the work is unknown.
Tokens per job, and per correct job once failed attempts bill too. The second bar is the one a budget should be built on.
Reasoning mode belongs in this section because it is the configuration flag that moves these numbers most. Thinking raised the score in 12 of 14 model pairs we tested, by up to 2.4× on the same weights, and no instruct configuration in the roster beats a 27B model with reasoning on, including a 753B one. Reasoning on is the setting a buyer should start from, and the case for switching it off is narrow, a non-critical function on infrastructure that cannot carry the extra tokens, checked first against the two exceptions, because the Qwen3.6 thinking builds score lower than their instruct builds and fabricate more than twice as often. The cost objection to reasoning is also weaker than it looks, since a correct task is what gets paid for, and on DeepSeek-V4-Flash the thinking build is cheaper per correct task than the instruct build because it fails less.
Fourteen base models scored in both modes. The two coral pairs are the exceptions, and both fabricate more with reasoning on.
2. Six model findings a capability leaderboard cannot show
Claude Opus 5 leads the public table at 4,558 on the PINNACLE Model Score, GPT-5.6 Sol and Claude Sonnet 5 sit 4.5% apart below it, and GLM-5.2 with reasoning on is the best open-weight configuration at 3,255, fourth overall. That ordering will look familiar. The rest of the table will not, because Signal65 PINNACLE scores the work enterprises deploy agents to do rather than general capability, rankings here will not match rankings elsewhere, and on this work the models separate on axes a capability leaderboard does not carry. The frontier closed models invent answers more often than the best open models do, one model family finishes the job in a fifth of the tokens anyone else needs, an open model is the right pick for a service desk, the entry-level Claude lands below an open 31B baseline, the best open weights are Chinese by a wide margin, and the open model most enterprises started with completes no whole job in our testing.
Fabrication is the failure enterprises ask about most and the one almost nobody quantifies, because scoring it requires knowing with certainty that an answer is absent, and that certainty is only available to whoever generated the corpus. At 128K context in our testing Claude Opus 5 produces an answer to 7.6% of questions the documents do not answer and GPT-5.6 Sol to 10.3%, while Qwen3.5-397B-A17B with reasoning on holds to 0.4%, GLM-5.2 to 1.5% and Qwen3.8-27B to 1.9%, and the low fabricators are not refusing by default, since they retrieve at 96 or better alongside. A retrieval miss announces itself and costs a delay. A fabrication arrives in the expected format in the expected place and flows into the next step, which is why the Customer Operations persona weights it above everything else and why an open model leads that persona ahead of both Claude models.
The smartest models make things up more often.
Inventing an answer, against finding the real one. The frontier closed models cluster to the right; the best open models hold under 2% and retrieve just as well.
Token efficiency separates one family from everyone. The three GPT-5.6 tiers land within 2% of each other at roughly 230 correct answers per million output tokens on the per-output basis in our testing, the next model is Claude Opus 5 at 108, and the best open configuration reads 92. GPT-5.6 Sol completes 97.9% of jobs in about 4,200 output tokens while the best open models need about 20,000 for 85 to 92%. Three tiers of very different size and price landing on the same efficiency constant reads either as a real advance in reasoning efficiency or as a technical limitation, a ceiling on reasoning tokens imposed to protect the infrastructure behind the API, and from the outside the two look the same, so we publish the number and the run settings and hold the interpretation. What the efficiency lead does not do is lower the bill, and the cost section below returns to it.
GPT-5.6 does the job in a fifth of the tokens. Efficiency, or a technical limitation?
The live results widget on its linear axis, per-output basis. 33 of 44 configurations crowd the left quarter of the chart, and the three GPT-5.6 tiers sit alone on the right. Explore it on the results page.
The persona layer turns one leaderboard into five answers, and the podium changes with the job. Claude Opus 5 leads Knowledge Worker, Data Analyst and Executive, GPT-5.6 Sol leads IT Professional by a margin inside the benchmark resolution over Opus 5, and Customer Operations belongs to GLM-5.2 with reasoning on, at 2,690 ahead of Claude Sonnet 5 at 2,504 and Claude Opus 5 at 2,305 in our testing. The persona weights fabrication resistance at 40% because a service agent that invents a policy is the failure that matters, GLM-5.2 fabricates on 1.5% of unanswerable questions against 3.9% and 7.6% for the two Claude models, and the recommendation carries its scope with it, since GLM-5.2 drops 18 points from governed to as-found data and is therefore the pick for an estate that has been organized. A buyer with a messy corpus should read that gap first, and the persona table is built so they can.
Five weightings of the same measurements, three different leaders. Personas never run separate scenarios.
The best model for customer operations agents is GLM-5.2. Not Claude. Not GPT.
Two hosted findings will draw the most mail. Claude Haiku 4.5 is priced as the entry point to the Claude family, and for agentic enterprise work the price is the only thing that is entry level, since it scores 638 against the open 31B baseline of 1,000, completes 28.9% of jobs overall and 9.2% on as-found data, gets 64.8% of individual deliverables right (the widest gap between doing the work and finishing it in the roster), and costs $1.09 per correct task on the full API bill, 80% of Claude Opus 5, because every failed attempt re-bills its context.
Claude Haiku 4.5 is cheap. It is not an agentic model.
The second finding is Llama. The open model the first wave of enterprise AI was built on, and possibly still the most widely deployed open model in production, completes no whole job in our testing, with Llama-4-Maverick and Llama-3.3-70B each finishing 0 of 280 jobs under the strict rule the whole benchmark uses while fabricating on 38% and 54% of unanswerable questions. Partial credit shows both models doing real intermediate work, which answers the harness objection, and a workflow that produces some correct outputs and never a complete set has still shipped nothing usable. Llama-3.3-70B was the reference model this program originally indexed at 1,000. On the current scale it scores 269.
Llama built the first wave of enterprise AI. It completes nothing.
The open-weight table has a shape general leaderboards do not show. The best open-weight configuration in our testing is Chinese, GLM-5.2 at 3,255, followed by DeepSeek-V4-Flash, Kimi-K3 and Qwen3.8-27B, all above 2,400. The best American open-weight configuration, Inkling-NVFP4, scores 1,003, sitting at the open baseline, and the best European entry, Mistral-Medium-3.5, scores 870. Six of the top ten configurations on the table are Chinese open-weight models, and the other four are American hosted models. The scale is reciprocal error, so the roughly 3× score gap reads as a weighted error budget about a third the size, or 91.1% job completion against 54.3%, and the boundary belongs in the same breath. This is enterprise agentic work specifically, 30 base models is a first cut, and there are more American open models still to test, the NVIDIA Nemotron family among them, so the table will be republished as models land. As of today, on this work, the open-weight race belongs to China and it is not close.
Today, the open-weight race belongs to China.
The roster, ranked. Hatched bars are hosted APIs; solid bars are open weights, colored by where they come from.
3. What we know about the silicon so far
The silicon lens is a deliberate work in progress. The capacity methodology was developed well before we engaged the GPU vendors, the launch effort put the model intelligence measurements first, and now that NVIDIA and AMD are on board we expect the fastest progress of the whole program here. Solid data exists today for three platforms, NVIDIA H200, NVIDIA B300 and AMD Instinct MI300X. Each is measured as a fully loaded 8-GPU node under the same agentic load on one harness, with every agent accumulating roughly 90K tokens of context as its job progresses. A rung counts only when every agent stays at or above 10 tokens per second and 95% of agents get a first token inside 60 seconds. The AMD Instinct MI355X results will follow on the system the lab began bringing online last week. Every capacity number here is 16-bit weights on vLLM on a single node, choices made for launch-window time and for putting the intelligence measurements first, and the expansion runs three ways from here, to FP8 and FP4 and derivatives like NVFP4 wherever a paired quality run shows the answers hold, to more serving engines and configurations, and to vendor-recommended setups run and disclosed as such. The serving stack is part of the configuration rather than a detail underneath it, since two of the eight engine configurations in the cross-engine study returned wrong answers at speed and a throughput-only measurement would have printed both as valid. For this first round every capacity comparison Signal65 PINNACLE publishes sits within one vendor, NVIDIA against NVIDIA and AMD against AMD, because until each vendor has supplied the configuration it would recommend to a customer running this workload, a cross-vendor chart reports our tuning choices alongside the silicon and a reader cannot tell which is which.
On the NVIDIA side the generational step is the largest we have measured. These are service-level multiples taken across a sweep of seven models, with every agent held at the floor and the admission gate enforced, rather than a peak read at a single favorable operating point, so they run lower than burst numbers and they describe capacity a buyer can bank. Across the seven models one B300 node holds 3.15× the concurrent agents of an H200 node at the same service level in our testing, from 1.86× on MiniMax-M3 to 6× on Qwen3.5-397B-A17B, where 16 agents becomes 96, and the three sparse Qwen mixture models gained 3.6× to 6× while the two dense models gained about 2.7×. A single average would wash out the extremes that matter most, the 6× gain on the largest model at one end and the 1.86× on MiniMax-M3 at the other, so the whole seven-model ladder publishes, failing rungs included.
NVIDIA B300 runs up to 6× the agents of H200 per node.
Seven models, one node each, same harness and gates on both boxes. The gain runs 1.86× to 6×, and no single number describes it.
The number that makes the step concrete is a headcount. One 8-GPU B300 node running Qwen3.5-35B-A3B sustains 320 concurrent agents, each holding about 90K tokens of context and each decoding at 18 tokens per second, with 95% of agents admitted inside 43 seconds. The rung above it at 328 agents fails the admission gate, so 320 is where the box tops out rather than where the sweep stopped. The same model holds 88 agents on H200, and the same B300 node holds 32 agents of the dense Gemma-4-31B. The model decides the headcount as much as the silicon does.
The live GPU vs GPU widget on Qwen3.5-35B-A3B. Filled points pass both gates, the open point above 320 fails one, and every rung publishes either way.
On the AMD side the first data set is one generation deep, and it already says two useful things. An 8-GPU Instinct MI300X node in our testing holds 128 concurrent agents of Qwen3.5-9B, and it holds 24 agents of the 397B-parameter Qwen3.5-397B-A17B, because 1.5 TB of HBM is why a model that size fits at all. The full seven-model ladder is in the chart below, a 6× spread on one box that mirrors the B300 story. Three of the seven cells ran AMD AITER FlashAttention kernels and the rest ran vLLM ROCm defaults, disclosed per cell, because on this platform the kernel choice moved capacity as much as anything else did. The second thing MI300X shows is that the box does not move the answers. The retrieval and knowledge scores measured on this hardware match the scores published everywhere else in this article, inside run-to-run noise, and the chart below carries the per-model numbers. Capacity varies by platform. Model quality does not, which is why quality is measured once per model, why a PINNACLE Model Score cannot be moved by choosing a box, and why the three lenses can share one measurement set. The published RIKER 2 study makes the same finding across NVIDIA, AMD and Intel silicon.
A 2023 GPU can run a 2026 agent workforce.
What one MI300X node holds, published within the AMD family for this first round.
The fairness result underneath the whole program. The same model scores the same on every box we ran it on, inside run-to-run noise.
Every platform number above carries its stack, and the table below is the disclosure behind them. The B300 system ran as NVIDIA delivered it. The H200 and MI300X systems were provisioned by the Signal65 AI Lab, and the MI355X column fills in when that system is characterized.
The test environment behind every capacity figure in this article. Full per-rung data publishes with the results.
4. Nobody buys tokens. The price of a correct answer.
Every AI price sheet quotes tokens, and no enterprise is looking to buy tokens for their own sake. It buys finished work, and until the token count is attached to a graded job the price of that work is unknown. Signal65 PINNACLE prices the work from the same runs that grade it. Hosted models carry their full API bill, input and cached input and output at published list rates, with the tokens spent on failed attempts billed to the tasks that finished. Open models are priced as an 8-GPU B300 node at $63 per hour, the median on-demand rate across GPU cloud providers researched 29 August 2026 in a market running $52.80 to $142.42, scaled by the share of the node-hour the work uses, which is how a fleet is bought. Across 37 priced configurations a correct task runs from under five cents to more than six dollars in our testing, and two findings follow that were not visible before. The bill for agentic work through an API is dominated by the input the agent re-reads on every round rather than the output it writes, 65% to 91% of the hosted invoice, so pricing hosted models on output tokens alone understated their cost per correct task by 3× to 11×. And a correct enterprise task from a frontier API costs more than a dollar once the whole bill is counted, $1.21 from GPT-5.6 Sol and $1.36 from Claude Opus 5, against the dime the sticker math implies. Cost per correct task never publishes alone, since the same view carries the PINNACLE Model Score on its other axis, so the reading order is to pick the intelligence band the job needs first and then buy that band at the lowest price per correct answer, and the frontier line in the widget below is there for exactly that decision.
Nobody knew what correct work costs. It is more than you thought.
The live cost widget at 50% node utilization rather than a best-case 100%. Up is more capable, left is cheaper per correct task, and the stepped line is the frontier. Explore it on the results page.
Priced that way, the choice between leasing hardware and buying tokens stops being an opinion and becomes a crossover point. DeepSeek-V4-Flash on a leased B300 node delivers a correct task for about 15 cents at full utilization in our testing, against 49 cents from GPT-5.6 Terra, 62 cents from Claude Sonnet 5, $1.21 from GPT-5.6 Sol and $1.36 from Claude Opus 5, and full utilization is a best case no deployment reaches, so the useful number is the crossover. The node stays cheaper than Terra down to 30% utilization and cheaper than Sonnet 5 down to 24%, Qwen3.8-27B with reasoning on does the job for about eight cents and undercuts Terra down to 16% while scoring higher, and buying the hardware outright rather than leasing it moves every one of those floors further down. The advantage is a mid-size phenomenon, real from 27B to about 300B parameters, and it disappears for the giant open models on a single 16-bit node, where Kimi-K3 costs more per correct task than Claude Sonnet 5 at any utilization because one node sustains only 24 agents of it. Sonnet 5 and Opus 5 also score higher than anything in the mid-size band, so the crossover prices the work, the score still picks the model, and both axes publish side by side.
Open models on your own GPUs: cheaper per correct task than every API.
The crossover, the number a lease-or-buy-tokens conversation should start with. Curves are the leased node; dashed lines are the full-bill APIs.
The hosted market itself has not settled, and that is the buyer’s leverage. The six hosted models run 7× apart on the PINNACLE Model Score and 18× apart on cost per correct task in our testing, and the two orderings do not match. Claude Sonnet 5 delivers 96% of the GPT-5.6 Sol score at 51% of its cost per correct task, because Sol bills the input an agent re-reads at $5 per million against $2 for Sonnet 5 and input is the bill. Claude Opus 5 charges 2.2× Sonnet 5 per correct task for 21% more score, and GPT-5.6 Terra charges 6.6× Luna for 1.6× the score, and the price ladder inside each family is steeper than its capability ladder. The most surprising value in the tier is GPT-5.6 Luna. At about seven cents per correct task on the full bill it is the only hosted model priced inside the self-hosted band, cheaper than 26 of the 33 configurations you could run yourself at full utilization, 1,358 correct tasks per $100 against 74 from Claude Opus 5, and it earns the price by writing the fewest output tokens in the roster and reading input at 20 cents per million. It scores 1,168, near the open baseline, and finishes about half of as-found jobs, so it is a pick for high-volume, supervised, lower-stakes work rather than for autonomy, and stated with both numbers it is the best correct-work-per-dollar buy among the closed APIs by a wide margin.
Closed APIs: 7× apart on capability, 18× apart on price, and not in that order.
Ranked twice, and the orderings do not match. The full API bill, list rates as of 28 August 2026.
The persona layer is where cost and capability meet the job. GPT-5.6 Sol is the top model for an IT Professional workload, where chained tool use decides the score, and the fifth pick for Customer Operations, where refusing to invent a policy decides it, a 4.4× swing between its best persona and its worst. Claude Opus 5 swings 3.4× in the same direction, while GLM-5.2 swings 1.4× and Qwen3.8-27B 1.3×, which makes the frontier closed models specialists and the best open models generalists. An enterprise standardizing on one model across many roles is safer with the generalist profile, because it has no persona where it collapses, and an enterprise deploying by role can give the specialist the one job it is best at, and the spread only exists because the five capabilities underneath are measured separately.
Six configurations, five persona scores each, scaled to each model’s own best. Spread is best persona over worst, a derived ratio of published persona scores.
The same table answers the constrained buyer with one name. Qwen3.8-27B with reasoning on is the best sub-50B configuration on every one of the five personas, fourth overall for Customer Operations, and the best retriever in the entire roster at 96.8 with a 1.9% fabrication rate. It is also one of the cheapest correct answers available, at about eight cents per correct task and 790 correct tasks per hour on one node, with the one weakness every mid-size model shares, a 24-point drop from governed to as-found data. Value per node-hour is a persona table too, and it crowns different models again, since Qwen3.8-27B leads Knowledge Worker, Customer Operations and Executive on expected value per hour on one B300 node while DeepSeek-V4-Flash leads Data Analyst and IT Professional, where deeper workflows reward its higher reliability, and neither is in the top three on the Model Score. Which model, and whether to buy it as tokens through an API or run it on GPUs you lease or own, has a different answer for each row of that table, and publishing the rows is the point.
What comes next
Signal65 PINNACLE is not a finished, frozen product, and it is not meant to become one. It is a living instrument, built to evolve as fast as the agentic ecosystem it measures, and our partners in the industry should be able to depend on it precisely because it keeps moving. The workloads will improve and change, and new personas and new use cases will join the five that launch today. The roster is a first cut, with the table republished as models land, the newest open-weight releases among them, and the AMD Instinct MI355X results publish when the new system is characterized. The precision work expands to FP8 and FP4 and derivatives like NVFP4 wherever a paired quality run shows the answers hold, and reduced precision was free in every pair we have tested so far. From there the coverage widens to more GPUs and platforms, more inference engines with more models measured across them, vendor-recommended configurations run and disclosed as such, energy and power measurement, and multi-node performance across solutions like GB300 NVL72, with the CPU side of the same agentic story already in early production. Cross-vendor capacity comparisons open once vendor configurations are in place across the board.
Signal65 PINNACLE also shows every layer of its work. The full concurrency sweep behind every capacity number is public, failing rungs included, the per-rung data is available to anyone who wants to check a figure, the governance framework sets out how a result is produced, reviewed, corrected and refreshed, and the methodology whitepaper explains every choice underneath these numbers. A model that finishes enterprise work reliably at a cost the work can carry is the better buy for that work, whatever the general leaderboards say, and Signal65 PINNACLE exists to make that case in numbers.
The leaderboard, the methodology whitepaper and the governance framework are live at stg-pinnaclesignal65-staging.kinsta.cloud. Results in this article reflect testing through 30 August 2026; API list rates researched 28 August 2026 and re-verified 29 August 2026.