Press Release

Nota AI Places 3rd in ICML 2026 Global Challenge, Validating Optimization Expertise on Open-Source AI Model 'Qwen'

July 13, 2026

Nota AI Places 3rd in ICML 2026 Global Challenge, Validating Optimization Expertise on Open-Source AI Model 'Qwen'
▶ Achieved 6.978x faster inference by combining quantization andspeculative decoding on a single NVIDIA A10G GPU
▶ Secured 3rd place and had two papers accepted at the ICMLAdaptFM workshop, whose organizing committee includes researchers from Amazonand Meta
▶ Hosted 'Nota AI - Korea Efficient Days,' strengthening tieswith OpenAI, Google, Qualcomm, and other global AI companies and researchers

Nota AI, a leading company in AI model compression and optimizationtechnology, placed 3rd in the Efficient Qwen Competition at ICML 2026, one ofthe world's premier machine learning conferences — validating its inferenceoptimization expertise on Qwen, an open-source AI model widely used bydevelopers across the global AI ecosystem.

ICML is a globally recognized conferencewhere leading tech companies, universities, and research institutions unveiltheir latest advances in machine learning and artificial intelligence. Thecompetition was held as part of the Resource-Adaptive Foundation ModelInference (AdaptFM) workshop, which focuses on technology for runninglarge-scale AI models efficiently under constrained computing resources, withan organizing committee that includes researchers from Amazon, Meta, and othermajor institutions.

The Efficient Qwen Competition challengedteams to run Qwen3.5-4B — an open-source large language model widely adopted byAI developers worldwide — on a single NVIDIA A10G GPU, and push responsegeneration speed as high as possible while preserving model performance. Inshort, the goal was to deliver the same quality of answer, faster. Against afield of roughly 40 teams from around the world, Nota AI achieved an averageinference speedup of 6.978x, earning 3rd place overall. The result carriesadded significance in that it validates Nota AI's optimization technology onQwen, a model actively used in real-world production environments, whilemaintaining answer quality without compromise.

Nota AI's approach paired its proprietaryquantization technology — which reduces a model's memory footprint andcomputational load — with speculative decoding, a technique that optimizes howAI generates responses. The company's quantization method is distinguished by afollow-up training step that compensates for performance loss, minimizingaccuracy degradation beyond what conventional approaches typically achieve.

Speculative decoding works by having alightweight draft model quickly propose candidate responses, which the fullmodel then verifies to produce the final output. Nota AI layered insliding-window attention on top of this process, prioritizing recent inputcontext to cut unnecessary computation and further sharpen inferenceefficiency.

Notably, the competition's top-performingteams converged on a similar formula: pairing quantization with speculativedecoding. That convergence points to a broader shift in real-world AIdeployment, where optimizing the inference process itself — not simplyshrinking model size — is becoming the defining competitive edge.

"This achievement validates ourinference optimization technology on Qwen, one of the most widely usedopen-source AI models in the global AI ecosystem," said Taeho Kim, CTO andCo-founder of Nota AI. "We will continue expanding the application of ouroptimization technology across a broader range of AI services, as well ason-device and edge AI environments."

Alongside the competition result, Nota AIhad two papers accepted at the ICML AdaptFM workshop, both addressingquantization techniques for Mixture-of-Experts (MoE) language modelarchitectures. The papers introduce optimization methods that take advantage ofthe MoE structure — which activates only a subset of expert sub-networks asneeded — to minimize performance loss even under tight memory and computebudgets.

This latest recognition builds on NotaAI's earlier win at an NVIDIA Nemotron hackathon, where the company took both atrack award and the overall grand prize with a data-driven MoE quantizationtechnique. Together, the ICML paper acceptances and the top-3 finish at theEfficient Qwen Competition affirm Nota AI's technical leadership in LLMoptimization — validated through both peer-reviewed research and real-worldperformance benchmarks.

During the ICML 2026 conference, Nota AIalso hosted 'Nota AI - Korea Efficient Days' near COEX in Samseong-dong, Seoul,introducing research trends and industry applications in efficient AI tocompany representatives from OpenAI, Google, Qualcomm, and other globalorganizations, along with visiting researchers and engineers. The event marksanother step in expanding the company's technical collaboration and networkacross the global AI ecosystem.

Nota AI continues to develop compressionand optimization technologies that run efficiently across diverse hardwareenvironments, extending its reach from on-device AI and physical AI intolarge-scale LLM inference optimization.

Related