Mon, 3 Aug 2026

Shifting from bigger is better to right-sized intelligence

The enterprise AI landscape in 2026 has matured far beyond the initial frenzy of experimentation. After years of testing massive foundation models, the conversation has fundamentally shifted from technological novelty to pragmatic, scalable deployment.

CIOs are no longer asking if they should adopt AI, but rather how to scale it responsibly, affordably, and repeatably across complex, multi-jurisdictional organisations.

Historically, North America has dictated the pace of technological adoption, with Europe and Asia following in its wake. However, the generative AI era has disrupted this traditional hierarchy. According to Ravi Saraogi, president and co-founder of Uniphore, the narrative has fundamentally shifted.

“Asia, especially Southeast Asia, is certainly holding up in terms of its own transition, almost at the same speed as US and European markets,” Saraogi observes. In fact, he notes that the region is arguably “a little ahead of the game compared to many other countries out there.”

This parity is not merely anecdotal; robust market data back it. The Asia-Pacific region has emerged as the fastest-growing segment for both Large Language Models (LLMs) and Small Language Models (SLMs). China leads with a projected 38.2% CAGR through 2036, fuelled by massive domestic investments expected to reach $38 billion by 2027. India follows closely at a 35.4% CAGR, driven by its IT services sector, while Southeast Asia boasts a 63% adoption rate of generative AI tools across at least one business function.

The driving force behind this regional surge is organisational maturity. As Saraogi succinctly puts it, “the whole challenge is not the readiness of technology in this situation in different parts of the world. It’s the readiness of the enterprise to adopt that technology change.”

The disruption of the contact centre and beyond

For Asian enterprises, this readiness is most visibly playing out in the front office. “The disruption of the contact centre with the use of AI is obviously taking centre stage again,” Saraogi notes. Legacy IT stacks in these environments are being leapfrogged by language models that act as the “brain” previously distributed across human executives.

By bundling decades of institutional knowledge into a single model, organisations can navigate complex customer conversations uniformly, bypassing the need for rigid, legacy-guided workflows.

Beyond the front office, back-office operations such as underwriting, claims processing, and KYC are being transformed.

“In the world of AI, where huge amounts of data can be utilised to train language models… can really solve a lot of this back-office operations-related alignment,” Saraogi explains. The goal is to adopt an “AI-first approach: how do you apply the technology first and then be led by people instead of people led by tech?”

The token economics trap: Why POCs fail to scale

Despite this readiness, a significant barrier remains: the treacherous journey from proof-of-concept (POC) to full-scale production. In the early days of generative AI, enterprises rushed to integrate massive, trillion-parameter foundation models to solve isolated use cases. While this approach suffices for low-volume pilots, it collapses under the weight of enterprise-scale economics.

Ravi Saraogi

“The moment you move to a full-scale deployment, and you turn around your usage from 1 to 100 times more, the amount of token cost and token economics goes out the window,” Saraogi warns.

“You are ROI-prohibitive day one, because the token consumption is so much that you cannot achieve an ROI.” Ravi Saraogi

This financial reality has catalysed an industry-wide pivot toward “right-sized intelligence.” According to HCLTech’s 2026 enterprise AI analysis, global AI spending approached $1.5 trillion in 2025, placing immense pressure on CIOs to demonstrate tangible business value. Consequently, the global SLM market, valued at $6.5 billion in 2024, is projected to skyrocket to $64 billion by 2034.

By distilling models to handle specific, high-frequency tasks, enterprises can reduce cost-per-inference by orders of magnitude, often running SLMs effectively on CPUs or edge hardware rather than relying on exorbitant hyperscale GPU footprints.

The sovereign AI mandate and regulatory labyrinths

Beyond economics, the imperative for data sovereignty is dictating enterprise AI architecture across the diverse APAC landscape. The region is a complex amalgamation of divergent regulatory regimes. China’s generative AI regulations mandate strict registration, content labelling, and accountability, with penalties reaching CNY 50 million. South Korea’s AI Basic Act, effective January 2026, established a comprehensive national framework, while India’s enforcement of the Digital Personal Data Protection Act introduced stringent obligations for AI systems processing personal data.

Navigating this labyrinth requires more than mere compliance; it demands architectural sovereignty. By 2027, 35% of countries are projected to rely on region-specific or sovereign AI platforms. “How do you create a sovereign AI architecture, which isn’t a single regulatory checkbox… It’s an architectural property that satisfies multiple regimes simultaneously,” Saraogi explains.

For enterprises operating on a hub-and-spoke model across Southeast Asia, relying on third-party foundational models poses an unacceptable risk. “Unless you own that IP, you are effectively putting up your entire ecosystem out there, which is at the whims and fancies of third-party dependency,” he cautions.

By fine-tuning domain-specific SLMs on private data, organisations maintain complete control over their inferencing, ensuring data never breaches sovereign boundaries.

Architecting the flywheel of intelligence

Achieving this sovereign, cost-efficient architecture requires solving the perennial enterprise headache: fragmented, siloed data. Historically, preparing data for AI use cases required armies of data scientists and months of manual engineering. Uniphore has approached this by deploying autonomous data agents that operate on a “zero copy, zero transport” methodology.

“What would have taken six to eight months now can be done in two days,” Saraogi notes, highlighting the platform’s ability to connect to over 400 backend ecosystems and create a unified data fabric without moving the underlying data.

This creates what Saraogi terms the “flywheel of intelligence.” It begins with data discovery and engineering, progressing to the autonomous creation of knowledge and context graphs. These graphs are then used to fine-tune domain-specific SLMs, which power autonomous agentic workflows.

Crucially, the system includes a trust layer—an evaluation framework that acts as a maker-checker on the agents’ outputs, feeding learnings back into the knowledge graph. This continuous, autonomous loop allows enterprises to approach Artificial General Intelligence (AGI) within their specific operational contexts, completely independent of forward-deployed engineers once the system is live.

The hybrid paradigm: Orchestrating LLMs and SLMs

The evolution of enterprise AI in 2026 is not a zero-sum game between LLMs and SLMs; rather, it is a strategic orchestration of both. The broader enterprise LLM market, valued at $7.57 billion in 2026, continues to serve as the foundational layer for complex reasoning, creative synthesis, and broad knowledge retrieval. General-purpose LLMs retain a 41.6% share of enterprise model deployments.

However, the operational heavy lifting is increasingly shifting to SLMs. The emerging paradigm is hybrid AI deployment: utilising LLMs as “orchestrators” for nuanced, multi-step queries, while delegating routine, latency-sensitive, and domain-specific operations to SLMs deployed at the edge.

This ensures that frontline environments—such as contact centres, clinical settings, and manufacturing floors—receive real-time, millisecond responsiveness without incurring the latency and costs associated with cloud-based monolithic models.

A smarter way to enterprise AI

As CIOs look toward the future, the external factors influencing their roadmaps are clear: absolute control over model learnings, strict adherence to divergent regional regulations, and uncompromising data sovereignty.

The era of relying on third-party, black-box foundation models is ending. By embracing sovereign architectures, right-sized SLMs, and autonomous data flywheels, Asia’s enterprises are not just participating in the AI revolution; they are defining its next, most pragmatic chapter.

Related:  The why and how of automating IT Ops

Related Stories

MORE STORIES

Subscribe