White Paper · Performance Engineering Series 06

Optimizing gpt-oss-120b Inference on the Tenstorrent TT-QuietBox

PDF · 19 pages

Cover Optimization and execution time breakdown Latency metrics comparison
CoverOptimization and execution time breakdownLatency metrics comparison

We ran gpt-oss-120b on Tenstorrent TT-QuietBox, identified bottlenecks through measurement, and optimized them. This paper walks through the process and the results.

Key results

−31.3% TPOT
Reduced time per output token (TPOT) from 144.85 ms to 99.49 ms.
+45.7% Decode speed
Increased per-user decode speed from 6.9 to 10.05 tokens/s.
−23.0% Mean latency
Reduced mean end-to-end latency (E2EL) per request from 25.0 s to 19.3 s.
3.68x Matmul speedup
Sped up MoE expert matrix multiplication, the main decode bottleneck, from 22.8 ms to 6.2 ms.

* Measured on TT-QuietBox (8 Wormhole chips, 96 GB GDDR6 in total) with one concurrent request and 128-token input and output.

The challenge this paper addresses

As hardware options for LLM inference expand, getting a large model to run does not always mean achieving practical response speeds. In initial testing on TT-QuietBox, gpt-oss-120b delivered a per-user decode speed of just 6.9 tokens/s. This paper explores how much performance can be gained from the available hardware by addressing underutilized compute cores, kernel launch overhead, and unnecessary data movement.

What sets this paper apart

  • Architecture-level bottleneck analysisExplains Tenstorrent’s Tensix cores and SRAM-centric architecture, and identifies why MoE matrix multiplications underutilize the available cores.
  • Parallelism and operator fusion in practiceShows how combining tensor parallelism (TP) with expert parallelism (EP) improves core utilization, while fusing operations into fewer kernels reduces overhead and data movement.
  • Implications for hardware selectionExamines the remaining challenges for coding agent workloads and discusses potential improvements with Galaxy systems and Blackhole-generation hardware, along with priorities for further testing.

Who this is for

  • AI infrastructure engineers evaluating on-premises LLM inference platforms, including alternatives to NVIDIA GPUs.
  • AI/ML engineers interested in large MoE model performance and practical inference optimization on Tenstorrent hardware.
  • Technical decision-makers selecting inference hardware based on measured performance and workload suitability.