See more industry reports and analysis at signal65.com

Benchmark Governance

How results are produced, reviewed, corrected, and refreshed.

Document Control
Version
1.0
Status
Draft
Issued
August 26, 2026
Last Revised
August 26, 2026
Applies to
PINNACLE v1.0 and every result published under it
Companion document
Signal65 PINNACLE Methodology Whitepaper v1.0
Version
Date
What changed
1.0
August 31, 2026
First published version (TBD)

Material changes to the rules in this document increment the version and are recorded above. Corrections to formatting, typography, and cross references do not. This document is versioned separately from the benchmark, and the PINNACLE version a result was produced under is published with that result.


Signal65 PINNACLE measures how well AI systems complete real enterprise work. Enterprises make purchasing decisions on the results, and silicon vendors and platform providers are compared by them. Both of those things impose an obligation, and this document is how Signal65 meets it.

Signal65 is an independent analyst and testing firm. What follows is the full set of rules that determine how a PINNACLE result comes to exist, who sees it before publication, what they can and cannot do about it, and what happens when something turns out to be wrong.

These rules are published in advance so they can be checked. They apply identically to every result, commissioned or not.

Principles


Five commitments sit underneath every specific rule in this document. Where a rule is silent or a situation is genuinely novel, these are what Signal65 decides against.

1. How a result is produced


Every published PINNACLE result comes from a run conducted by Signal65 on the PICARD execution harness. Signal65 configures the environment, executes the run, and reads the output. Results submitted by a third party are not published.

1.1 What is disclosed with every result
Disclosed
Detail
Hardware
Platform, accelerator or processor model, node configuration, and the deployment shape measured
Software
Serving framework and version, driver and runtime versions, and any relevant library versions
Model
Model identifier, revision, precision, and the sampling parameters used
Test scope
Which scenarios were run, sample counts, and the service level thresholds applied
Tuning
Any configuration tuning applied, and who supplied it
Provenance
Whether each figure is measured, derived, or estimated, and the transform where a figure is derived
Funding
Whether the run was commissioned and by whom, or self initiated by Signal65
1.2 Tuning supplied by vendors

Vendors may supply recommended configurations and Signal65 will use them. This produces a better result for the vendor and a more representative result for a buyer, because it reflects what the vendor would actually deploy. Supplied tuning is disclosed alongside the result with its source named.

Signal65 does not accept configuration changes that are specific to the benchmark rather than to production. If a setting would not be recommended to a customer deploying the same workload, it does not belong in a PINNACLE run.

1.3 Serving framework and scope

A platform result reflects the serving framework it was measured on. A different framework can produce a different result and in some cases can reorder platforms. Signal65 states the framework and version with every cross platform comparison and treats broader framework coverage as an expansion of the benchmark rather than an implied property of it.

The same applies to deployment shape. A single node measurement describes a real deployment but not every deployment, and the shape measured is always stated.

2. Vendor review


Any vendor whose product appears in a result is offered a review before publication. The purpose of that review is factual accuracy and nothing else.

2.1 What review covers
2.2 What review does not cover
A vendor has no approval right and no veto over publication.

Review does not extend to whether a result is published, when it is published, how it is characterized, or what it is compared against. A vendor who disagrees with a finding but cannot identify a factual error is welcome to say so publicly, and Signal65 will publish a result over a vendor objection.

2.3 The review window

The standard review window is five business days from the point a result is shared. If a vendor does not respond within the window, publication proceeds. If a vendor identifies a factual error, the correction is made and, where it materially changes a figure, the affected test is re-run before publication.

Five business days is the window for hardware and platform results, which is what most PINNACLE work consists of at launch. Those results turn on configuration detail that takes a vendor real time to check, and a week does not erode what the result is worth to a reader.

New model coverage compresses the window to one business day. A model result is only useful while the release is current, and holding a model that shipped yesterday for five days means publishing late on the thing a reader wants first. Compression is workable because a new model is measured on a platform Signal65 has already characterized and disclosed, so a short window asks a reviewer for a check rather than an investigation. The deadline is stated when the result is shared.

The length of the window is the only thing that changes between the two. Everything in 2.1 and 2.2 applies identically at either length, and a factual error identified inside a compressed window is corrected on the same terms as any other. Signal65 does not delay publication on request. Both windows are fixed so that review cannot function as a mechanism for controlling timing.

3. Funding and independence


Signal65 publishes two kinds of PINNACLE result and labels every one of them.

Type
What it means
Self initiated
Signal65 chose the subject, funded the work, and published regardless of the outcome. Broad model coverage and platform generational comparisons are typically self initiated
Commissioned
A customer paid for the testing. The methodology, scoring, and disclosure rules are identical to a self initiated run, and the result is published whatever it shows
3.1 What commissioning buys and does not buy

A commissioned engagement buys Signal65 time, hardware access, and a report. It does not buy influence over the grade, favorable framing, the right to suppress a result, or preferential treatment in how a result is compared to others.

Where a commissioned run produces a result unfavorable to the customer, the customer may decline to promote it. They cannot prevent Signal65 from publishing it if it forms part of a comparison Signal65 has published.

3.2 Access

Hardware access is a practical necessity for this work and Signal65 accepts it from vendors. Access does not confer any right over methodology, timing, or the content of a result, and any vendor supplied hardware used to produce a published figure is disclosed with that figure.

4. Corrections


Errors will happen.

4.1 How a correction is made
4.2 What triggers a correction
4.3 Raising an issue

Anyone may raise a suspected error, not only vendors and not only customers. Signal65 investigates and responds, and where an issue is substantiated the correction follows the process above.

5. Re-running and refresh


A benchmark result describes a stack at a moment in time. Software moves quickly in this field, and a figure that is not refreshed becomes a claim about history rather than about a product.

5.1 What triggers a re-run
Trigger
Response
A serving framework or runtime version materially changes performance
The affected results are re-measured and the prior figures superseded
A vendor supplies a corrected or improved configuration
The affected results are re-measured with the new configuration disclosed
A methodology error is identified
Every result the error touched is re-measured
A new model release
Covered on its own timeline rather than by re-running the existing roster
A vendor requests a re-run
Considered on the same basis as any other request, and granted where there is a substantive reason
Signal65 does not commit to a fixed refresh calendar. A published cadence that is missed is worse than an honest description of what causes a re-run, and the pace at which this field moves makes any fixed schedule a promise that will break.
5.2 New models

Signal65 aims to cover significant new model releases across the platforms available in the lab as close to release as practical. Coverage depends on hardware availability and on whether the model runs on the platforms in question, and both are disclosed where a model is absent from a comparison.

6. Versioning and reference points


6.1 Benchmark versions

PINNACLE is versioned. A change that alters scenarios, scoring, or the composition of a metric is a version change, and results from different versions are not directly comparable. Version changes are announced with a description of what changed and what it means for prior results.

A change that adds coverage without altering how existing tests are scored is not a version change. New models, new platforms, and new scenarios added alongside existing ones do not break comparability.

6.2 The reference points

PINNACLE uses two normalization constants. A reference model anchors the quality scale, and a reference platform anchors the speed scale. Both are named in the methodology whitepaper.

The reference platform is chosen for capability rather than preference. It is the fastest platform available to Signal65 that can host the full model roster, because a reference that cannot run every model under test cannot anchor a comparison across them. The selection reason is stated wherever the reference appears.

6.3 When a reference moves

A reference point moves when it can no longer do its job. For the reference model, that is when it falls far enough behind the field that the scale compresses at the top. For the reference platform, that is when it can no longer host the models being tested, or when a materially more capable platform becomes available on the same terms.

A reference change is a version change. Prior results are restated against the new reference where the underlying data allows it, and where it does not, prior results are marked as belonging to the earlier scale rather than being silently carried forward.

7. What Signal65 does not do


Stated explicitly, because a governance document that only describes commitments leaves the interesting questions unanswered.

8. Scope and limitations


Governance covers how results are produced. It does not make the benchmark measure more than it measures, and the limits are documented in full in the methodology whitepaper. The most important are stated here because a reader of this document should not have to go looking.

Raising an issue


Questions about methodology, suspected errors, and requests for re-measurement can be raised with Signal65 directly. Substantiated issues are corrected under section 4 regardless of who raises them.

The full methodology, including scenario construction, scoring composition, and the measured basis for every published metric, is documented in the Signal65 PINNACLE methodology whitepaper.