See more industry reports and analysis at signal65.com

New models and the AMD Instinct MI355X join the PINNACLE GPU results.

Three models enter as initial results, the MI355X arrives at up to 6.5x the agents of MI300X in our testing, and the results views read across vendors for the first time, generation to generation.

Signal65 · Ryan Shrout · Data basis: results through September 14th, 2026


The PINNACLE GPU results views now cover twelve models on four GPU platforms. The AMD Instinct MI355X is on the board for the first time, with seven models measured against MI300X, and three models enter as initial results, Kimi-K3, GLM-5.2 and an MXFP8 build of MiniMax-M3, tested so far on the B300 and the MI355X. Every result comes from the same PINNACLE agent harness, a fully loaded 8-GPU node running a roughly 90K-token agentic workload at rising concurrency, and every result carries its serving environment in the results views.

Two numbers describe each platform. The headline is the number of concurrent agents a node sustains at the service level, every agent decoding at 10 tokens per second or better and 95% of agents receiving a first token inside 60 seconds. The modifier is speed, the decode rate each agent sees at that ceiling and at the lighter agent loads both platforms were measured at, because two platforms can tie on the headline while one is the more responsive system, and a platform can trail on headcount while leading at the loads most deployments run at.

All four platforms run at their best measured serving configuration to date, but none of them is finished. Serving stacks are still being tuned in our lab, AMD and NVIDIA engineers are working on configurations for these workloads, and the MI355X is the newest system with room to move, so every number here is a point in time and the results views will update as that work lands.

MI355X over MI300X

The step from MI300X to MI355X is a median 2x on agents at the service level across the seven models measured on both in our testing, up to 6.5x on MiniMax-M3, and the MI355X is also the faster platform at every agent count the two share. At the load where MI300X reaches its ceiling, the MI355X delivers about 2x the tokens on Qwen3.5-35B-A3B, 1.8x on Qwen3.5-397B-A17B and 1.6x on Qwen3.5-122B-A10B. The gain is largest on mixture-of-experts models, 2x to 2.4x against 1.3x to 1.8x on the dense models, which is consistent with the step from 192 GB to 288 GB per GPU on workloads that carry roughly 90K tokens of context per agent.

Agents at the service level, MI355X against MI300X

Figure 1. Concurrent agents at the service level on one 8-GPU node, MI355X against MI300X, seven models. Multiples at right are MI355X over MI300X.

Model MI355X agents · tok/s MI300X agents · tok/s Agents Speed at MI300X's ceiling
Qwen3.5-9B 168 · 1,914 128 · 1,255 1.3x 1.5x at 128 agents
Qwen3.5-27B (dense) 60 · 590 40 · 463 1.5x 1.2x at 40
Gemma-4-31B (dense) 36 · 361 20 · 279 1.8x 1.1x at 16
Qwen3.5-35B-A3B 152 · 1,534 64 · 640 2.4x 2.0x at 64
Qwen3.5-122B-A10B 96 · 925 48 · 501 2.0x 1.6x at 48
Qwen3.5-397B-A17B 56 · 542 24 · 270 2.3x 1.8x at 24
MiniMax-M3 52 · 508 8 · 77 6.5x 1.9x at 8

Table 1. Agents at the service level and aggregate decode at that point, with the MI355X's aggregate multiple at the agent count where MI300X reaches its ceiling (nearest shared rung for Gemma-4-31B). In our testing.

Every one of these comparisons runs at the same precision on both generations, so none of the MI355X's native lower-precision formats contributes yet, and the MI300X figures include attention configurations tuned on that platform over the summer while the MI355X results use mostly default configurations. With that software work still ahead, we expect the generational improvement to grow rather than shrink, and the MiniMax-M3 multiple in particular is a figure we expect to revisit as an MI300X re-test lands.

Across vendors, generation to generation

For the first time the results views show AMD and NVIDIA on the same models. The comparisons are generation-matched, H200 against MI300X and B300 against MI355X, on one harness, one service level and one fully loaded 8-GPU node. In our testing the prior generation was a split decision, MI300X ahead on four of nine models and H200 on four with one tie, and the current generation is a B300 lead on five of eight, with the MI355X ahead on Kimi-K3 and Gemma-4-31B and the two tied on MiniMax-M3.

PINNACLE Agent Harness

Figure 2. Agents at the service level on one 8-GPU node. Left, the prior generation, H200 against MI300X. Right, the current generation, B300 against MI355X. Kimi-K3 and GLM-5.2 are initial results.

Both generational steps are large. NVIDIA's is larger on headcount, a geometric mean of 3.15x for B300 over H200 across seven models against 2.2x for MI355X over MI300X, and AMD's carries a property the NVIDIA step does not, faster at every shared rung. The B300's ceilings on the original pairs are mostly set by the 60-second admission gate, while every MI355X ceiling is a decode-floor crossing with first-token latency still measured in seconds, which is the first place the two vectors diverge. Neither number is final.

Qwen3.5-122B-A10B and Qwen3.5-35B-A3B

On the two mid-size Qwen3.5 models the B300 lead is widest on both vectors. Qwen3.5-122B-A10B sustains 192 agents on the B300 against 96 on the MI355X in our testing, and at 88 agents the B300 delivers about 1.8x the tokens and 1.8x the decode rate per agent. Qwen3.5-35B-A3B is the small model that behaves like a fleet, 320 agents at the service level on one B300 node against 152 on the MI355X, and at 160 agents the B300 is still delivering about 25 tokens per second to every agent where the MI355X has crossed its floor.

Qwen3.5-35B-A3B - per-agent decode rate as concurrency rises

Figure 3. Qwen3.5-35B-A3B, tokens per second per agent as concurrency rises on all four platforms. The largest marker on each curve is the platform's agent count at the service level; open markers fail a gate.

Model B300 agents · tok/s MI355X agents · tok/s Agents Speed at matched load
Qwen3.5-122B-A10B 192 · 1,964 96 · 925 B300 2.0x 1.8x at 88 agents (1,607 against 887 tok/s)
Qwen3.5-35B-A3B 320 · 5,845 152 · 1,534 B300 2.1x 2.7x per agent at 160 (25.3 against 9.3 tok/s)

Table 2. Agents at the service level with aggregate decode at that point, and the B300 multiple at a load both platforms were measured at. In our testing.

That fleet is a lookup and extraction workforce rather than an autonomous one, since the model completes about 31% of KAMI jobs strictly while scoring 95.1 on RIKER retrieval, and the B300 pays for the 320 with admission latency, a 37.6-second median first token at its ceiling against 2.7 seconds on the MI355X. The B300 sweeps on both models have no rung below 88 and 120 agents, so the light-load comparison is a gap the next update will address.

MiniMax-M3, the same 52 agents for two different reasons

On MiniMax-M3 the B300 and the MI355X sustain the same 52 agents at the service level in our testing, and they arrive there by different roads. The B300 is 3.7x faster per agent at 8 agents, 2.4x at 16 and 1.4x at the shared ceiling, its aggregate peaks at 48 agents, and the rung past the ceiling is a cliff, with per-agent decode halving and first-token latency climbing past 9 seconds at 64 agents. The MI355X starts lower and declines gently, with first-token latency flat at 0.2 seconds through 56 agents and aggregate still rising at the ceiling.

MiniMax-M3 - per-agent decode rate as concurrency rises

Figure 4. MiniMax-M3, tokens per second per agent as concurrency rises. Dotted multiples are B300 over MI355X at the same agent count.

Agents B300 tok/s · per agent · first token MI355X tok/s · per agent · first token B300 multiple
8 536 · 67.0 · 0.1 s 146 · 18.3 · 0.1 s 3.7x
16 609 · 38.0 · 0.1 s 254 · 15.9 · 0.1 s 2.4x
32 736 · 23.0 · 0.1 s 393 · 12.3 · 0.2 s 1.9x
48 804 · 16.8 · 0.2 s (B300 peak) 479 · 10.0 · 0.2 s 1.7x
52 729 · 14.0 · 0.5 s (ceiling) 508 · 9.8 · 0.2 s (ceiling) 1.4x
56 462 · 8.3 · 1.0 s (fails the floor) 504 · 9.0 · 0.2 s (fails the floor)

Table 3. The shared rungs for MiniMax-M3, aggregate tokens per second, tokens per second per agent and median first-token latency. In our testing.

For a buyer that reads as the same headcount with a more responsive agent on the B300, and for AMD it reads as competitive performance with the B300 on this model at the ceiling, with the B300 the faster platform at lighter loads.

MiniMax-M3 at MXFP8

Reduced precision is where the serving stack shows most clearly. On the B300 the MXFP8 build of MiniMax-M3 sustains at least 72 agents against 52 for the 16-bit weights in our testing, with peak aggregate decode within 2%, so the halved weight footprint converts into agents. On the H200 the same build sustains fewer agents than 16-bit, 20 against 28, and on the MI355X an initial run on an earlier system held 32 against 52, on a coarse sweep that passed at 32 and failed at 64 with nothing between. Both current-generation platforms carry native 8-bit hardware paths, so the difference reads as how far each serving stack had taken them at the time of the run, and the MI355X MXFP8 rerun on the current system is queued.

MiniMax-M3 - 16-bit weights against the MXFP8 build, agents at the service level

Figure 5. MiniMax-M3, agents at the service level for the 16-bit weights and the MXFP8 build on each platform. The B300 MXFP8 sweep stopped at 72 before finding the ceiling.

Platform 16-bit agents · tok/s MXFP8 agents · tok/s Change in agents
B300 52 · 729 at least 72 · 713 at least +38%
H200 28 · 527 20 · 392 −29%
MI355X 52 · 508 32 · 367 (initial, earlier system) −38% on a coarse sweep

Table 4. The MXFP8 checkpoint is 444 GB against 854 GB for the BF16 weights, which frees roughly 50 GB per GPU for KV cache on a 288 GB platform. In our testing.

Kimi-K3 and GLM-5.2, initial results

Kimi-K3 is the largest model in the set, served from its 4-bit MXFP4 weights on both platforms, and the one place the current-generation comparison runs the other way on both vectors. In our testing the MI355X sustains 32 agents at the service level against 20 on the B300, and from 16 agents onward it is also the faster platform per agent, 20.4 tokens per second against 15.1 at 16 agents, with the B300 holding a small edge at the lightest shared load, 31.2 against 26.4 at 8. The B300 result came on a pre-release vLLM build, since no tagged release supported the model at the time, so this is as much a snapshot of serving-software maturity on a new model as a silicon result, and we expect the NVIDIA number to move.

GLM-5.2 runs the other way. In our testing the B300 sustains 28 agents against 12 on the MI355X and is about 2.2x faster per agent at 4 agents, a lead on both vectors. The MI355X result is four rungs on a serving build whose support for the model was days old, and its aggregate falls between 8 and 12 agents in a way the silicon does not explain, so we expect this number to move more than any other here once AMD's configuration work reaches it.

Per-agent decode rate

Figure 6. Initial results, tokens per second per agent as concurrency rises. Left, Kimi-K3. Right, GLM-5.2. B300 and MI355X only. The largest marker on each curve is the platform's agent count at the service level; open markers fail a gate.

Model B300 agents · tok/s MI355X agents · tok/s Ahead Note
Kimi-K3 (MXFP4) 20 · 217 32 · 326 MI355X 1.6x MI355X faster per agent from 16 agents up (20.4 against 15.1 tok/s at 16); B300 on a pre-release vLLM build
GLM-5.2 (BF16) 28 · 520 12 · 121 B300 2.3x B300 2.2x faster per agent at 4 agents; MI355X result is four rungs on a days-old serving build

Table 5. The two new models measured on both current-generation platforms, initial results. In our testing.

Where this leaves things

Our read of these results is that the B300 still performs like the leader on some of the most recent models. Its margins are widest on the two mid-size Qwen3.5 models and on GLM-5.2, and at the lighter agent loads both platforms were measured at it is the faster platform per agent on most of the models, with Kimi-K3 and Gemma-4-31B the exceptions. The MI355X reads as a strong competitor rather than a distant second. It matches the B300's agent count on MiniMax-M3, leads on Kimi-K3 and Gemma-4-31B, roughly doubles its predecessor on agents at the service level, and it is the platform whose results we expect to move the most as serving stacks and vendor configuration work catch up to the silicon.

None of this accounts for cost. Performance per dollar, performance per watt and the price of the node underneath each result are the questions a buyer asks next, and they are the subject of a later PINNACLE post rather than this one. For now, the sentences we are comfortable with are that on agents at the service level the current generation is a B300 lead with real exceptions, that on speed per agent it is a B300 lead with the same two exceptions, and that both sentences carry a date on them.

More is coming

More is coming soon. Reduced-precision builds (FP8, MXFP8 and NVFP4) will run alongside the 16-bit results wherever a paired quality run shows the answers hold, the serving-engine comparison that today covers SGLang and vLLM on H200 and MI300X extends to MI355X and B300 with TensorRT-LLM added where a supported path exists, multi-node results on GB300 are in early testing and look promising, the MI355X model set expands to match B300 coverage, every sweep gains a common ladder of agent counts so matched-load comparisons exist at light load as well as at the ceiling, and the cost side of these results, performance per dollar and per watt on each platform, gets a post of its own.

Thanks

These results exist because AMD and NVIDIA both put engineering time and hardware access behind an independent benchmark that publishes unconditionally, and both teams are working through serving configurations with the Signal65 lab on their current systems. We are grateful for the partnership and look forward to pushing PINNACLE further with both. Every number in this post is live in the PINNACLE results views at pinnacle.signal65.com, with the environment behind each result in its tooltip.