Why Your AI’s Workflow Matters More Than Its Model

Why Your AI’s Workflow Matters More Than Its Model — featured image

by

The AI Model Is Not the Whole Automation

When you test an AI tool for your business, it is natural to compare model names, benchmark scores and advertised capabilities. However, a recent Nvidia research result points to a more practical lesson: the software “harness” surrounding an AI model can determine whether the system completes useful work or gets stuck, repeats mistakes and makes risky decisions.

For a Malaysian SME, this matters because most business tasks are not single questions. You may need an AI system to read an enquiry, check stock, confirm delivery areas, prepare a quotation, request approval, update your customer relationship management system and remind a staff member to follow up. That is a chain of actions, not just a chat response.

The model provides reasoning and language capabilities. The harness provides the operating discipline: which tools the AI can use, what information it can remember, what rules it must follow, how it receives feedback and when a human must take over.

What Happened

According to TechCrunch, Nvidia researchers tested Claude Opus 5 on ARC-AGI-3, an interactive reasoning benchmark made up of 2D games with no instructions. Without a customised harness, the model scored 30%. With a harness designed to manage memory and include a supervisor component, it achieved a 100% score in the reported test.

A harness is the software wrapper around an AI model. It can include tools, memory controls, rules, runtime functions and feedback mechanisms. Nvidia’s researchers used a system called Agentic Variation Operators, or AVO, which included a supervising agent that could prompt the main agent when it became stuck, repeated an unproductive path or moved towards a dead end.

The idea is similar to having a manager review a project while an employee performs the detailed work. The main AI handles the task, while the supervisor checks direction and encourages a better next step. Nvidia executive Adel El Hallak described the agent as more than a model: it includes the model, the scaffolding, tools, runtime and associated skills, as reported by TechCrunch.

The result does not mean model quality is irrelevant. A stronger model may understand instructions better or handle more complex information. It does mean that selecting a model is only one part of designing a dependable AI workflow.

Why This Matters for Malaysian SMEs

Your business probably operates across several disconnected tools. Customer enquiries may arrive through WhatsApp, Instagram, Facebook, email or a website form. Product details may sit in a spreadsheet. Delivery information may be in a courier portal. Invoices may be prepared in accounting software, while approvals happen in a staff group chat.

An AI model placed on top of this setup may write a convincing reply, but it cannot reliably complete the whole process unless its harness defines what it can access and what it must do. For example, an AI sales assistant should not merely draft “Your order is available.” It should check the current stock record, confirm the selected variation, identify whether the delivery address is within your service area and send the enquiry to a human if any information is missing.

Consider a Malaysian catering business handling corporate lunch orders. A suitable harness could guide the AI through this sequence:

Stage What the AI should do Control required
Enquiry Extract date, location, guest count and dietary needs Ask for missing information
Availability Check the internal booking calendar Do not confirm without a valid slot
Quotation Use approved menu and service rules Require approval for unusual requests
Follow-up Schedule a reminder if the customer has not replied Stop after a defined number of reminders

The same principle applies to workshops, wholesalers, tuition centres, professional services firms and online sellers. If your AI must perform multiple steps, the workflow around it deserves as much attention as the prompt or model selection.

It is also relevant to data protection and operational safety. The TechCrunch report notes that long-running AI systems have been observed deleting files or databases and taking unsafe actions while pursuing objectives. Your business may not be running an advanced research agent, but the risk principle is clear: an automated system should have limited permissions, approval checkpoints, activity logs and an easy way to stop execution.

Before asking, “Which AI model should we use?”, ask, “What should the AI be allowed to do, what information should it see, and who checks its work?”

How to Build a Better AI Harness for Your Business

You do not need to create a complex research system to apply this lesson. Start by mapping one repetitive process from beginning to end. Choose a task where the inputs are reasonably consistent and where the result can be checked.

  • Define the outcome: State exactly what completed work means, such as a confirmed appointment, an approved quotation or a properly categorised lead.
  • Separate steps: Break the process into information gathering, checking, drafting, approval and follow-up.
  • Connect only necessary tools: Give the AI access to the relevant database, calendar or document folder, rather than your entire business system.
  • Create memory rules: Specify what can be remembered, how long it is retained and which information must be verified again.
  • Add a supervisor: Use a review step that checks for missing fields, contradictions, repeated attempts or unusual requests.
  • Set human handovers: Escalate complaints, uncertain information, sensitive records and exceptions to a named staff member.
  • Keep an audit trail: Record the request, actions taken, tools used and final outcome so you can investigate errors.

For example, an AI receptionist for a small clinic should not independently answer every health-related question. Its harness can handle appointment availability and basic administrative details, while directing medical questions, urgent situations and ambiguous messages to trained staff. The business benefit comes from reducing repetitive administration without pretending that every conversation is suitable for automation.

The Bigger Picture

The AI market is moving from isolated chatbots towards agents that can carry out longer sequences of work. That shift changes how you should evaluate technology. A polished demonstration may show an impressive answer, but your business needs consistent performance across real customer messages, incomplete data, staff changes and unexpected exceptions.

Research cited by TechCrunch also suggests that the harness can affect operational efficiency. Databricks CEO Ali Ghodsi reportedly said that using the same model with different harnesses could significantly change usage and operating requirements, with a poor harness potentially doubling the level of model use. The practical message is to test the complete workflow, not just the model in isolation.

Nvidia’s wider position is that open harnesses and open components can give users more control over accuracy, infrastructure, runtime behaviour and security. For an SME owner, that does not mean you must build everything internally. It means you should understand where your data goes, which permissions the AI has, whether workflows can be changed and whether you can export your records if you switch platforms.

When reviewing an AI automation project, ask the vendor to demonstrate failure handling. What happens when a customer gives an incomplete address? What if stock data is outdated? Does the system ask for clarification, pause for approval or confidently continue? Can you see every action? Can a staff member stop the process? These answers are often more important than the model’s headline score.

Your Practical Next Step

Choose one workflow that takes your team at least several hours each week. Document its current steps, common exceptions and approval points. Then design a small harness around it: clear instructions, limited tools, controlled memory, a supervisor check and human escalation. Test it with real examples, including messy messages and missing information.

The strongest AI setup for your business may not be the one with the most impressive model name. It may be the one that remembers the right details, follows your process, knows its limits and brings you into the decision at the right moment. That is where reliable automation is built: not around a model alone, but around the operating system that helps the model work safely for you.

Ready to Streamline Your Operations?

Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →