White Paper · Performance Engineering Series 06
Optimizing gpt-oss-120b Inference on the Tenstorrent TT-QuietBox
PDF · 19 pages
We ran gpt-oss-120b on Tenstorrent TT-QuietBox, identified bottlenecks through measurement, and optimized them. This paper walks through the process and the results.
Key results
- −31.3% TPOT
- Reduced time per output token (TPOT) from 144.85 ms to 99.49 ms.
- +45.7% Decode speed
- Increased per-user decode speed from 6.9 to 10.05 tokens/s.
- −23.0% Mean latency
- Reduced mean end-to-end latency (E2EL) per request from 25.0 s to 19.3 s.
- 3.68x Matmul speedup
- Sped up MoE expert matrix multiplication, the main decode bottleneck, from 22.8 ms to 6.2 ms.
* Measured on TT-QuietBox (8 Wormhole chips, 96 GB GDDR6 in total) with one concurrent request and 128-token input and output.
The challenge this paper addresses
As hardware options for LLM inference expand, getting a large model to run does not always mean achieving practical response speeds. In initial testing on TT-QuietBox, gpt-oss-120b delivered a per-user decode speed of just 6.9 tokens/s. This paper explores how much performance can be gained from the available hardware by addressing underutilized compute cores, kernel launch overhead, and unnecessary data movement.
What sets this paper apart
- Architecture-level bottleneck analysisExplains Tenstorrent’s Tensix cores and SRAM-centric architecture, and identifies why MoE matrix multiplications underutilize the available cores.
- Parallelism and operator fusion in practiceShows how combining tensor parallelism (TP) with expert parallelism (EP) improves core utilization, while fusing operations into fewer kernels reduces overhead and data movement.
- Implications for hardware selectionExamines the remaining challenges for coding agent workloads and discusses potential improvements with Galaxy systems and Blackhole-generation hardware, along with priorities for further testing.
Who this is for
- AI infrastructure engineers evaluating on-premises LLM inference platforms, including alternatives to NVIDIA GPUs.
- AI/ML engineers interested in large MoE model performance and practical inference optimization on Tenstorrent hardware.
- Technical decision-makers selecting inference hardware based on measured performance and workload suitability.