Press Release
July 7, 2026

Nota AI, a pioneering leader in AI modelhardware-aware optimization and compression technology, announced today that itis enabling enterprise customers to maximize the performance of their AI modelson Amazon Web Services (AWS) dedicated AI chips. Holding status as a ValidatedAWS Partner, Nota AI offers specialized compression and tuning servicesdesigned to unlock the peak performance of customer workloads on AWS customsilicon environments, including AWS Trainium.
The newly launched service targetsenterprise clients facing a shortage of specialized AI optimization talent. Bytuning models precisely to AWS’s high-performance AI chips, Nota AI helpsbusinesses achieve significant cost efficiencies and boosted inference speeds.Nota AI conducts rigorous technical diagnostics and tuning tailored to eachclient's specific model architecture. This allows enterprises to fullycapitalize on the superior price-performance and energy efficiency that AWScustom silicon brings to production-grade workloads.
Specifically, Nota AI’s new serviceaccommodates enterprises evaluating or adopting AWS Trainium—purpose-built byAWS for high-performance deep learning training with optimal tokeneconomics—and AWS Inferentia, engineered for low-cost, high-throughputinference.
The service is powered by NetsPresso®, NotaAI’s proprietary hardware-aware AI optimization platform that automates theentire lifecycle from model compression to target hardware deployment.NetsPresso® can shrink AI model sizes by up to 90% or more while maintainingstrict baseline accuracy, delivering multi-hardware optimization within asingle, streamlined workflow. The newly rolled-out service follows a structuredthree-stage roadmap: Proof of Concept (PoC) & Diagnostics, Model Porting& Compression, and Target Performance Tuning.
AWS Trainium and Inferentia have alreadyestablished a proven track record among global industry leaders. Leadingcompanies such as Anthropic, Apple, Databricks, Uber, Ricoh, and Decartleverage AWS custom silicon. Notably, as of Q3 2025, AWS's Trainium-relatedbusiness experienced explosive growth, surging 150% quarter-over-quarter into amulti-billion dollar segment. Nota AI’s service focuses on eliminatingtechnical friction, allowing businesses to integrate and deploy their customworkloads onto these powerful, market-proven AWS AI chips rapidly.
This service launch stands on thefoundation of Nota AI’s deep, proven engineering expertise. The company hasaccumulated extensive experience tuning Large Language Models (LLMs) nativelywithin AWS Trainium and Inferentia environments. In a recent notable benchmark,Nota AI applied its compression technology to a high-performance32-billion-parameter language model, successfully reducing the model size by68% while keeping the accuracy loss below 1%. The company also holds a vastportfolio of porting and tuning diverse model architectures, allowing clientsto tap into the capabilities of AWS custom silicon instantly.
"AWS AI chips offer unparalleledprice-performance, and Nota AI is uniquely positioned to help enterprises fullyunlock that performance potential for their custom models through experttuning," said Myungsu Chae, CEO of Nota AI. "We are committed todelivering tangible cost-reduction outcomes and performance breakthroughs tocorporate clients seeking to transition their infrastructure to AWS AIsilicon."
Nota AI has consistently validated itsmarket-leading technology through strategic partnerships across a diversehardware ecosystem, collaborating with Samsung Electronics in mobile, Arm insemiconductor IP design, FuriosaAI in data center accelerators, and Mobilint inedge devices. The rollout of this AWS-centric optimization service marks asignificant expansion of Nota AI’s technology footprint into the cloudinfrastructure domain. Moving forward, Nota AI intends to solidify its positionas the ultimate model compression and tuning partner across every environmentwhere AI workloads run.