Why Your SME Needs a 10M-Token AI That Never Leaves Your Office

Why Your SME Needs a 10M-Token AI That Never Leaves Your Office — featured image

by

The AI That Reads Your Entire Company’s History — Without Cloud Leaks

Imagine an AI assistant that could review every email, contract, customer complaint, and operations log your business has produced in the last decade — all at once, in a single conversation, without sending a single byte to a foreign server. That sounds like science fiction for a Malaysian SME running on a tight IT budget. But a new model called Pokee-Isaac 28B claims to make it real: a 10-million-token context window that runs entirely inside your own boundary — be it your office server, a Malaysian data centre VPC, or even a laptop-class device.

For Malaysian business owners, this isn’t just another AI lab milestone. It’s a direct answer to the question you’ve probably been asking since ChatGPT became a workplace tool: “How do I use AI on my confidential data without risking a PDPA violation?” As of the latest amendments, Malaysia’s Personal Data Protection Act requires strict consent and security for personal data. Sending customer data to overseas cloud AI endpoints is legally murky at best. This new class of in-boundary AI models may be the practical escape hatch you’ve been waiting for.

What Happened: A 28-Billion-Parameter Model That Holds 10 Million Tokens

Pokee AI released Pokee-Isaac 28B, a text-only foundation model with a 10-million-token context window, specifically designed to be deployed inside a customer’s own infrastructure — not accessed via a public cloud API. According to the announcement, the research team reports a 93.3% score on RULER at 10M tokens, meaning the model can retrieve and reason over information buried in roughly 7,500 pages of text in one go. Until now, such long-context capability was “available almost exclusively from cloud endpoints,” which excludes regulated industries and on-device applications where data cannot leave the boundary at all.

The model is not open-weight; it’s licensed for deployment inside a VPC, on-premises, or on-device, with Day-0 support for vLLM and SGLang inference engines. The company advertises single-GPU serving starting from an RTX 4090 or equivalent, though the published measurements only come from a single B200-class GPU, so treat the consumer-GPU claim as vendor guidance rather than a reported result. On agentic benchmarks, Isaac scores 70.94 on BFCL v4 (vs. GPT-5.6 Luna’s 70.61), averages 0.662 on τ³-bench (ahead of Gemini 3.5 Flash Lite’s 0.631), and records the lowest attack success rate on the DTAP red-teaming benchmark at 35.6% combined — all while maintaining 82.5% benign task success.

The efficiency numbers are striking: under the RULER workload on one B200 GPU, time-to-first-token is 23.6 seconds at 1M context and 72.9 seconds at 10M, while prefill throughput rises from 42,400 to 137,200 tokens per second. The pricing is listed as $0.15 and $1.00 per million input/output tokens (provisional), but that’s for the hosted API — the licensed in-boundary deployment is what makes this story interesting for your business. It also runs fully on-device on Intel Arc Pro B70, Core Ultra Series 3, and Qualcomm Snapdragon X2 Elite.

“When enough usable context is available in-boundary, memory hierarchies and compression become optional rather than required.” — Pokee AI research paper on Pokee-Isaac 28B

Why This Matters for Malaysian SMEs: Your Data, Your Rules

Let’s be honest: you probably don’t need a 10-million-token context window. Your SME’s entire annual sales conversation history might be 500,000 tokens. But you do need the principle it enables. Malaysian SME owners often face a false choice: use powerful cloud AI and risk exposing customer or employee data to third-party processors, or stay “safe” and miss out on AI productivity gains. Pokee-Isaac’s existence demonstrates that the AI industry is waking up to a third option: bring the model to your data, not your data to the model.

Concrete use cases for a Malaysian SME: a legal firm reviewing multi-year tenancy agreements and employment contracts in one pass, flagging renewal dates and liability clauses without uploading client files to a foreign server. A logistics company doing incident forensics over full archive logs — think of a customs audit or a dispute with a supplier — where the entire evidence trail fits in one prompt. A medical clinic or health-tech startup handling patient records under the Private Healthcare Facilities and Services Act; these records cannot leave the clinic’s boundary without patient consent. With an in-boundary model, you can run AI-assisted triage on historical patient intake forms while keeping everything behind your own firewall.

Even for a retail SME, consider customer support: instead of a chatbot that vaguely answers from a knowledge base, an in-boundary agent could read every past chat, every return request, and every product description across your entire history — then respond with full context. Since it runs on your hardware (a decent GPU workstation or a rented VM in a Malaysian cloud region), you control exactly who has access. That’s a different risk profile than sending customer conversations to a third-party API where the provider’s terms govern data use.

The Bigger Picture: The End of Context Pruning for On-Prem AI

This release signals a shift in what “enterprise AI” means. Long-horizon agents — AI that performs multi-step tasks like “audit all vendor invoices over the past 5 years and flag discrepancies” — have been held back by context limits. Previous models had to summarize or compress intermediate steps, losing fidelity. Pokee-Isaac’s 10M-token context, combined with in-boundary deployment, means that summarization and context pruning become optional. For Malaysian SMEs adopting automation, this is a quiet revolution: you no longer need to architect complex “memory hierarchies” or data pipelines just to make an agent remember what it did five minutes ago.

But be cautious. The model’s weights are not published, so you’re dependent on Pokee AI’s license. That’s a vendor lock-in consideration. And the single-GPU consumer claim (RTX 4090) hasn’t been independently verified — only B200-class measurements are published. For a small business, renting a cloud VM with a single B200 or A100 GPU in Singapore/Malaysia might be the realistic entry point, though on-device support for Intel and Qualcomm chips means future laptops could run this locally.

The deeper message: Malaysian SMEs no longer have to choose between AI capability and data sovereignty. As more models follow this “inside the customer boundary” pattern, your automation strategy can include AI that reads everything and leaks nothing. That’s a competitive advantage your larger competitors may not be moving on yet.

Key Points to Remember

  • 10M-token context: reads ~7,500 pages in one shot (93.3% RULER at 10M) — source: Pokee AI release coverage
  • In-boundary deployment: VPC, on-prem, or on-device; not open-weight, licensed model
  • Agentic performance: leads BFCL v4 at 70.94, averages 0.662 on τ³-bench; one benchmark (Terminal-Bench 2.1) lost to a cloud baseline
  • Security: lowest DTAP attack success rate (35.6% combined) among tested models
  • Hardware: runs on B200 (verified) and claims RTX 4090 (unverified); also on Intel and Qualcomm chips

For your Malaysian SME, the action item is simple: audit which of your current AI tools are processing confidential data externally. If you’re already using a hosted AI for internal documents, start exploring licensed, in-boundary models like Pokee-Isaac 28B the next time you upgrade your automation stack. The future of SME AI isn’t just smarter models — it’s models that respect your boundaries.

Ready to Streamline Your Operations?

Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →