Aolani, a Singapore-founded neocloud, has announced the launch of the Aolani Token Factory, a managed inference platform designed to allow organizations to deploy and scale AI models on a pay-per-token basis without managing underlying GPU infrastructure. This move positions Aolani as the first Singapore-founded neocloud to offer production-grade managed inference at scale, addressing the growing demand for efficient AI deployment solutions.
The platform responds to the accelerating need for production-grade inference infrastructure as global AI companies expand operations in Singapore and enterprises worldwide invest in AI for business outcomes. The Token Factory aims to close the accessibility gap between AI experimentation and production-scale deployment by offering a highly compliant and high-performance path for AI-native companies and enterprises.
With per-token metering, customers can pre-purchase credits and pay based on token consumption, avoiding capital-intensive GPU investments. Aolani manages the entire inference stack, including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization, enabling customers to scale consumption without provisioning additional infrastructure.
At launch, the platform supports leading open-source models such as DeepSeek, GLM, Kimi, and Qwen, with plans to expand the catalog based on customer demand. Customers can also deploy their own models using OpenAI-compatible APIs. For enterprise clients with strict compliance and data residency needs, dedicated capacity and data isolation options are available.
The Token Factory focuses on three core production use cases: AI agents for high-volume inference and workflow automation, enterprise AI applications like internal copilots and knowledge assistants, and coding agents for code generation and review. Sea Xu, Applied AI Research Lead at Aolani, emphasized the platform's high-performance inference stack and the goal to keep pace with Southeast Asia's rapidly evolving AI ecosystem.
Nicholas Chia, CEO of Aolani, highlighted the flexibility and cost-effectiveness of the service: "Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure."
The launch of the Aolani Token Factory signifies a shift in how organizations access AI compute, potentially lowering barriers for startups and enterprises to adopt advanced AI capabilities. By eliminating the need for upfront GPU investments, the platform could accelerate innovation and democratize access to AI in Asia and beyond. Interested parties can register interest at Aolani's Token Factory page.
