Press Release

Nota AI Unveils Compressed Version of Moonshot AI's Kimi K3, Cutting GPU Requirements From Eight to as Few as Four

August 6, 2026

Nota AI Unveils Compressed Version of Moonshot AI's Kimi K3, Cutting GPU Requirements From Eight to as Few as Four
- Reduces NVIDIA B300 GPU requirements from eight to six or four, improving inference efficiency and infrastructure utilization
- Following its compressed Upstage Solar Open 2 model, Nota AI proves its compression technology on a frontier-scale model with 2.8 trillion parameters
- Proprietary "Non-uniform Expert Pruning" technique applied to a massive MoE model that is difficult to quantize further

Nota AI (CEO Myungsu Chae), a company specializing in AI model compression and optimization, announced today that it has compressed Moonshot AI's frontier-scale model Kimi K3, cutting the number of graphics processing units (GPUs) required to run it by as much as half. The achievement demonstrates Nota AI's proprietary technology on a frontier-class AI model with 2.8 trillion parameters.

Kimi K3 is a frontier-scale, open-weight AI model that uses a Mixture-of-Experts (MoE) architecture, selectively activating only the modules needed for a given input to improve computational efficiency. Its high performance comes at a cost: running the model requires eight of NVIDIA's latest high-performance B300 GPUs, creating a significant cost and infrastructure burden for companies and developers looking to adopt it.

To make frontier-scale models more accessible, it is necessary to reduce model size and computational load while preserving core performance. However, Kimi K3 was designed from the training stage with weights and activations quantized down to 4-bit and 8-bit precision, respectively — leaving limited room for further compression through conventional quantization methods, which reduce size and computation by lowering numerical precision.

Nota AI addressed this challenge with its proprietary "Non-uniform Expert Pruning" technology. The method analyzes the characteristics of each layer and the importance of each module, preserving key experts while selectively removing modules that contribute the least. Unlike conventional approaches that prune every layer at the same rate, Nota AI's technique applies different pruning ratios by layer, minimizing performance degradation.

Using this approach, Nota AI released two compressed versions of Kimi K3, with expert modules reduced by 25% and 50%, respectively. The number of NVIDIA B300 GPUs required to run the model was reduced accordingly, from eight to six and four. This allows companies and developers to operate large-scale AI models with substantially lower infrastructure costs.

Performance was also preserved. In Nota AI's internal evaluation, its 50% compressed model — which applies different pruning ratios by layer — outperformed REAP (Router-weighted Expert Activation Pruning), a method that prunes experts at a uniform rate across all layers, on most benchmarks. Notably, the model scored 2.52 points higher on GPQA-Diamond, which evaluates advanced scientific reasoning, and 2.22 points higher on IFEval, which measures a model's ability to accurately follow instructions.

The company said the results demonstrate the applicability of its technology to large-scale Mixture-of-Experts models. Nota AI's earlier compressed version of Upstage's Solar Open 2 model, released on Hugging Face, surpassed 68,000 cumulative downloads within two weeks of release, drawing strong interest from the market. Nota AI said the response reflects growing demand from companies and developers seeking to run high-performance AI models with fewer computational resources and lower costs.

"Building on our proprietary Non-uniform Expert Pruning technology, we successfully compressed the frontier-scale Kimi K3 model following Solar Open 2, cutting the GPUs required by as much as half while preserving core performance and improving the computational efficiency of large-scale AI models," said Kim Taeho, Chief Technology Officer (CTO) of Nota AI. "We will continue expanding the scope of our technology so that a wide range of frontier-scale AI models can be used efficiently even with limited computational resources."

The compressed versions of Solar Open 2 and Kimi K3 released by Nota AI are available on the global AI platform Hugging Face.

Related