Why this AI story matters to your business
AI tools are becoming easier for small businesses to adopt, but a new enterprise survey highlights a warning you should not ignore: an AI system can pass internal testing and still create a customer-facing problem after launch.
According to VentureBeat’s July 2026 VB Pulse research, 49% of surveyed organisations had experienced an AI agent or large language model feature passing company testing before disappointing customers. Nearly a quarter, or 24%, said this had happened more than once.
For a Malaysian SME, this is not only an issue for large technology companies. The same pattern can appear when you use an AI chatbot for customer enquiries, an automated system for quotation preparation, a tool that updates stock records, or software that drafts messages to customers. A system may look reliable during a small test but behave differently when it meets real customers, unusual requests, incomplete information, or local business conditions.
What Happened
VentureBeat surveyed 108 people from organisations with at least 100 employees. The publication clearly described the findings as directional rather than a complete market census because the respondents were self-selected and the sample size was limited. Even so, the results show an important tension between growing confidence in automated evaluation and continuing evidence that testing can miss real-world failures.
In July, 13% of respondents said they completely trusted automated evaluation, up from 5% in the previous month. At the same time, the number identifying poor alignment between tests and real-world results as their biggest concern fell from 29% to 19%, according to the survey findings.
However, the proportion reporting at least one customer-facing incident after an AI feature passed internal testing remained almost unchanged: 50% in June and 49% in July across 265 enterprise responses over the two months. This does not mean that 49% of AI interactions fail. It means that almost half of the surveyed organisations had experienced at least one such incident in the previous year.
The most surprising result concerned deployment decisions. Overall, 67% of respondents either already allowed an agent to push code or change a system without human approval in selected low-risk situations, or were preparing to do so within the following year. Among organisations that had already experienced a customer-facing AI failure after testing, 85% were pursuing this no-approval approach, compared with 61% among organisations reporting no comparable incident.
A successful test is evidence that a system performed well under selected conditions. It is not proof that the system will always behave correctly in live business operations.
Why This Matters for Malaysian SMEs
You may not be running hundreds of AI agents, but your business can still be exposed to the same weakness. Consider a Kuala Lumpur retailer using AI to answer WhatsApp enquiries. During testing, the chatbot may correctly respond to common questions about opening hours and delivery areas. In production, a customer may ask about a damaged parcel, a combined promotion, or a product variation that was not included in the test data. A confident but incorrect reply can quickly become a service complaint.
The risk is also relevant to service firms, wholesalers, clinics, tuition centres, restaurants, contractors and professional practices. An AI tool that drafts quotations may omit an item. An automated stock assistant may rely on outdated inventory data. A scheduling system may confirm an appointment without checking staff availability. A document assistant may create a polished reply that misinterprets a customer’s instructions. The output can appear professional while still being wrong.
Local operating conditions make practical monitoring especially important. Your staff may communicate in English, Bahasa Malaysia, Mandarin, Tamil or mixed language. Customers may use abbreviations, informal spelling and voice messages. Your business may also depend on several separate systems for accounting, customer records, inventory, delivery and payment status. An AI workflow that works in one system may produce an error when information is missing or inconsistent across systems.
This does not mean you should avoid automation. It means you should decide where human review is necessary. For example, your chatbot can answer routine product questions automatically, while complaints, refunds, credit decisions, contract changes and unusual delivery requests must be sent to a staff member.
The Bigger Picture
The survey separates two activities that are often treated as the same: pre-launch evaluation and live production monitoring. Testing asks whether an AI system appears ready before release. Monitoring asks whether it continues to produce accurate, appropriate and safe results after customers begin using it.
The July research found that only 26% of valid respondents used inline quality assertions, such as automated checks or guardrails that review live output quality. Another 26% mainly tracked transaction traces, including system activity and inputs and outputs, while 24% focused on gateway measures such as latency, errors and usage indicators, according to VentureBeat Intelligence’s report.
For your SME, speed and system availability are useful indicators, but they do not confirm that an answer is correct. A chatbot can respond quickly, show no technical error and still give a customer the wrong delivery date. A workflow can complete successfully while writing an incorrect value into your records. Your monitoring should therefore include business-quality checks, not only technical performance.
A practical AI control plan for your SME
| Business area | Safer automation approach | Human review trigger |
|---|---|---|
| Customer enquiries | Automate common questions using approved information | Complaints, refunds, legal threats or unusual requests |
| Quotations | Let AI prepare a draft from your product and service records | Final pricing, discounts, tax details and special terms |
| Inventory | Use AI to identify low-stock patterns and prepare alerts | Any automatic stock adjustment or purchase commitment |
| Marketing | Generate campaign drafts and customer segments | Claims, regulated statements and final customer lists |
| Administration | Extract information from documents into a review queue | Banking, payroll, contracts and compliance documents |
What you should do next
Start by listing every AI-enabled feature your team uses, including tools inside software subscriptions. Record what information each tool can read, what action it can take and who is accountable when something goes wrong. This simple inventory often reveals that an AI tool has more access than your team realised.
Next, create a small test set based on real Malaysian customer behaviour. Include mixed-language questions, incomplete details, spelling mistakes, delivery exceptions, returns, public holidays and requests that require escalation. Do not test only the easiest examples. Your goal is to discover where the system needs boundaries.
Keep an approval step for high-impact actions. AI can draft, classify, summarise and recommend, but you should be cautious about allowing it to issue final refunds, alter customer records, approve credit, change contractual terms or send sensitive information without review.
Finally, review live results every week during the early rollout. Track customer corrections, escalations, repeated questions and cases where staff had to repair an AI-generated action. Use these incidents to update your instructions, knowledge base and escalation rules. If an error reaches a customer, treat it as a process signal rather than simply blaming the tool.
The lesson from the survey is not that automation is unsuitable for SMEs. The lesson is that deployment should not be the finish line. Your best results will come from combining automation for repetitive work with clear limits, business-focused checks and human attention where the consequences matter most.
Ready to Streamline Your Operations?
Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →