Press Release

Nota AI’s Optimization Research Earns Global Recognition Two Papers Accepted at EMNLP 2026, a Leading NLP Conference

August 25, 2026

Nota AI’s Optimization Research Earns Global Recognition Two Papers Accepted at EMNLP 2026, a Leading NLP Conference

- One paper accepted to the EMNLP 2026 Main Conference and another to Findings, both introducing new techniques for MoE quantization
- The EMNLP acceptances follow two paper acceptances at an ICML 2026 workshop in July, further validating Nota AI’s technology within the global research community
- Optimization technology cuts the number of GPUs required to run Qwen3.8-Max from 24 to four, improving efficiency for LLMs and data centers

Nota AI, an AI model optimization technology company, is strengthening its position in AI optimization as its research continues to earn recognition from the global academic community.

Nota AI announced on Aug. 25 that two of its papers on large-scale AI model optimization had been accepted at EMNLP 2026, a leading international conference in natural language processing (NLP), with one selected for the Main Conference and the other for Findings. Roughly 18,000 papers were submitted this year, with the Main Conference accepting just 15.4% of submissions.

The two papers address performance degradation that can occur when models based on a Mixture-of-Experts (MoE) architecture are quantized. MoE architectures activate only a subset of relevant experts for each input, but the full model must still be stored in memory, requiring substantial GPU and memory resources. Quantization can reduce this burden, but small numerical shifts can change which experts are selected and degrade model performance.

The paper accepted to the Main Conference introduces “MENDS-MoE,” a methodology that accounts for how quantization affects expert selection at subsequent steps and preserves the ranking of experts near the selection boundary. In 4-bit and 3-bit quantization experiments across three MoE models, the method achieved higher average accuracy and stronger language-model performance than comparison techniques in most evaluation settings.

The Findings paper proposes “OPERA,” a methodology that concentrates optimization on changes in expert selection that affect a model’s final output rather than uniformly correcting every shift. The two studies complement one another by helping preserve routing decisions after quantization while focusing optimization on the changes that matter most to response quality.

Earlier this year, two of Nota AI’s MoE quantization papers were also accepted at the “Resource-Adaptive Foundation Model Inference” (AdaptFM) workshop at ICML 2026.

Nota AI is applying its research-validated optimization methodologies to large language models and data center AI infrastructure. Most recently, its optimization technology enabled Qwen3.8-Max, a model with more than 1 trillion parameters, to run on four NVIDIA B300 GPUs, compared with 24 GPUs for the original model configuration. The company also enabled Moonshot AI’s Kimi K3 to run on as few as four NVIDIA B300 GPUs, compared with eight, and Upstage’s Solar Open 2 to run on two NVIDIA H100 GPUs, compared with eight.

“As AI models grow to hundreds of billions or trillions of parameters, the industry’s competitive focus is shifting from raw model performance to operational efficiency,” said Myungsu Chae, CEO of Nota AI. “This achievement demonstrates that years of AI optimization research have given Nota AI the technical foundation needed to improve the efficiency of large language models and data center AI infrastructure. We will continue investing in R&D to keep pace with architectural changes in state-of-the-art AI models and expand the application of our core optimization technology to make AI infrastructure more efficient.”

Related