Press Release

Nota AI Places 3rd in ICML 2026 Global Challenge, Validating Optimization Expertise on Open-Source AI Model 'Qwen'

July 13, 2026

Nota AI Places 3rd in ICML 2026 Global Challenge, Validating Optimization Expertise on Open-Source AI Model 'Qwen'
- Achieved 6.978x faster inference by combining quantization andspeculative decoding on a single NVIDIA A10G GPU
- Secured 3rd place and had two papers accepted at the ICMLAdaptFM workshop, whose organizing committee includes researchers from Amazonand Meta
- Hosted 'Nota AI - Korea Efficient Days,' strengthening tieswith OpenAI, Google, Qualcomm, and other global AI companies and researchers

Nota AI, a leading company in AI model compression and optimization technology, placed 3rd in the Efficient Qwen Competition at ICML 2026, one of the world's premier machine learning conferences — validating its inference optimization expertise on Qwen, an open-source AI model widely used by developers across the global AI ecosystem.

ICML is a globally recognized conference where leading tech companies, universities, and research institutions unveil their latest advances in machine learning and artificial intelligence. The competition was held as part of the Resource-Adaptive Foundation Model Inference (AdaptFM) workshop, which focuses on technology for running large-scale AI models efficiently under constrained computing resources, with an organizing committee that includes researchers from Amazon, Meta, and other major institutions.

The Efficient Qwen Competition challenged teams to run Qwen3.5-4B — an open-source large language model widely adopted by AI developers worldwide — on a single NVIDIA A10G GPU, and push response generation speed as high as possible while preserving model performance. In short, the goal was to deliver the same quality of answer, faster. Against afield of roughly 40 teams from around the world, Nota AI achieved an average inference speedup of 6.978x, earning 3rd place overall. The result carries added significance in that it validates Nota AI's optimization technology on Qwen, a model actively used in real-world production environments, while maintaining answer quality without compromise.

Nota AI's approach paired its proprietary quantization technology — which reduces a model's memory footprint and computational load — with speculative decoding, a technique that optimizes how AI generates responses. The company's quantization method is distinguished by a follow-up training step that compensates for performance loss, minimizing accuracy degradation beyond what conventional approaches typically achieve.

Speculative decoding works by having a light weight draft model quickly propose candidate responses, which the full model then verifies to produce the final output. Nota AI layered in sliding-window attention on top of this process, prioritizing recent input context to cut unnecessary computation and further sharpen inference efficiency.

Notably, the competition's top-performing teams converged on a similar formula: pairing quantization with speculative decoding. That convergence points to a broader shift in real-world AI deployment, where optimizing the inference process itself — not simply shrinking model size — is becoming the defining competitive edge.

"This achievement validates our inference optimization technology on Qwen, one of the most widely used open-source AI models in the global AI ecosystem," said Tae-ho Kim, CTO and Co-founder of Nota AI. "We will continue expanding the application of our optimization technology across a broader range of AI services, as well as on-device and edge AI environments."

Alongside the competition result, Nota AI had two papers accepted at the ICML AdaptFM workshop, both addressing quantization techniques for Mixture-of-Experts (MoE) language model architectures. The papers introduce optimization methods that take advantage of the MoE structure — which activates only a subset of expert sub-networks as needed — to minimize performance loss even under tight memory and compute budgets.

This latest recognition builds on Nota AI's earlier win at an NVIDIA Nemotron hackathon, where the company took both a track award and the overall grand prize with a data-driven MoE quantization technique. Together, the ICML paper acceptances and the top-3 finish at the Efficient Qwen Competition affirm Nota AI's technical leadership in LLM optimization — validated through both peer-reviewed research and real-world performance benchmarks.

During the ICML 2026 conference, Nota AI also hosted 'Nota AI - Korea Efficient Days' near COEX in Samseong-dong, Seoul,introducing research trends and industry applications in efficient AI to company representatives from Open AI, Google, Qualcomm, and other global organizations, along with visiting researchers and engineers. The event marks another step in expanding the company's technical collaboration and network across the global AI ecosystem.

Nota AI continues to develop compression and optimization technologies that run efficiently across diverse hardware environments, extending its reach from on-device AI and physical AI into large-scale LLM inference optimization.

Related