White Papers - GPU Optimization and AI Performance

Bringing the Large-Scale MoE Model GLM-5.1 to Practical Speeds on a Single NVIDIA H200 Node

Making the 754B-total-parameter MoE model GLM-5.1 practical on a single NVIDIA H200 node. A hands-on look at how KV cache optimization and other techniques cut P90 TTFT to roughly 1/15.

Case Study: Optimizing Ore Blending with Quantum-Inspired Technology

Optimizing Copper Ore Blending with Fixstars Amplify’s Quantum-Inspired Annealing Technology

Empirical Comparison of NVIDIA H100 and B200

Explore the performance leap of the Blackwell architecture. This report demonstrates how a single NVIDIA B200 node delivers up to 4× higher cost-effectiveness for LLM pre-training compared to H100 configurations.

Practices in Large-Scale Distributed GPU Configurations

Benchmarking performance on a 64× NVIDIA H100 cluster. Learn how to achieve 3x better cost-performance for massive model inference and optimize 8-node distributed training environments.

Visualizing and Improving GPU Utilization in Autonomous Driving AI Development

A record of technical collaboration with Sony Honda Mobility, using AIBooster to visualize and improve GPU utilization in their autonomous driving AI development environment

NVIDIA H200: AI Acceleration and Performance Engineering in Practice

This white paper explores proven methods to harness the full power of the latest GPUs and accelerate real-world AI workloads.

Revolutionizing AI Development Efficiency Through GPU Optimization with Fixstars AIBooster

A hands-on guide to overcoming the hidden GPU-utilization challenges most companies miss, with actionable strategies to boost AI efficiency.