AI Agents That Code: Why Malaysian SMEs Should Watch

AI Agents That Code: Why Malaysian SMEs Should Watch — featured image

by

An AI just built your app. Can you trust it?

You’re running a business in Malaysia. You probably don’t write code, and you shouldn’t have to. But the systems you rely on — your customer database, your booking system, the app you’re planning to launch — are increasingly built by AI. Not “assisted by AI.” Built line by line by an AI tool.

That raises a practical question: can you trust it? Not in a philosophical sense. Will it get the rules right that stop one customer from seeing another customer’s private data? Will it hold up at 9am on a busy Monday? This week, Supabase — the company behind a widely used database platform — released an open-source benchmark called Supabase Evals that offers one of the most honest answers yet. Source

Here’s what you need to know, and why it matters for how you choose vendors, buy automation, and protect your customers’ data.

TL;DR

Supabase open-sourced a benchmark that tests AI coding agents on real software tasks — building databases, debugging failures, fixing security rules — and publishes the scores. Source The strongest agents now pass 100% of build tasks with no help. Smaller models catch up when given detailed written guidance. The biggest remaining weaknesses are the things that matter most to you: security rules and database structure.

What this means, in plain language

An AI coding agent is a tool like Claude Code, OpenAI’s Codex, or OpenCode. You describe a task in plain English — “build a customer signup form that saves to a database” — and the agent writes the code, runs it, watches for problems, and fixes its own mistakes. Think of it as hiring a junior developer who works at remarkable speed and sometimes needs the same instruction twice.

A benchmark is a standardised test. Supabase built its test tasks from real problems their customers actually encountered — support tickets, bug reports, GitHub issues. Source These are not textbook exercises. They’re the messy, frustrating situations that come up in real businesses: a failed function, a broken security rule, a database that needs repair.

Here’s the part that makes this different: the agents work against a real, running system, not a simulation. They get genuine tools, a genuine environment, and one retry before being graded. Scoring combines automated checks with an AI judge. Source The results are public at supabase.com/evals.

One term worth understanding: RLS, or Row-Level Security. It’s the rule system that decides who can see what inside a database. In business terms: the rules that stop Customer A from seeing Customer B’s order history. Get this wrong, and you don’t have a technical bug — you have a privacy breach.

How this applies to Malaysian SMEs

It changes how you evaluate vendors. You may not write code, but you hire people who do — web developers, automation companies, app agencies. These vendors increasingly rely on AI to build faster. That’s good, but “faster” only helps if the outcome is correct. Supabase Evals gives you a concrete way to understand what AI tools have actually been tested on. Next time an agency tells you their AI “handles the backend,” you can ask: tested against what? Source

Your data security depends on this. Malaysia’s Personal Data Protection Act holds you responsible for the personal data you collect — not your vendor, and certainly not an AI agent. The benchmark’s findings show that even the best agents struggle most with security rules and database migrations. Source Agents were caught hand-writing database migrations instead of using safer, standardised approaches, and verifying authentication by hand instead of using proper security packages. Source For a Malaysian SME, the takeaway is clear: if an AI writes the rules that protect your customers’ data, a human who understands the stakes must review them.

It shows how to get better results from AI. The benchmark revealed something encouraging: when agents were given “skills” — written guidance on how to work properly — weaker models jumped from 78% to 100%. Source The lesson for your business is that the quality of an AI setup isn’t just about the model. It’s about how well someone has documented your industry, your security requirements, and your existing systems. An automation vendor who invests in teaching their AI about your business will get different results from one who just types in a prompt and hopes.

It maps directly to business automation. If you use automation tools to manage appointments, invoices, or customer records, you depend on software that touches real data. The tasks Supabase tests — building a database structure, debugging a function, fixing a security policy — are the invisible plumbing underneath every automation project. When that plumbing fails, it’s rarely the “automation” part at fault. It’s the database, the security rules, or the integration. This benchmark is the first time we can measure, publicly, exactly how good AI has become at that plumbing.

“A benchmark is only as good as the tasks it tests. Supabase built theirs from real support tickets and bug reports — which means passing means something in the real world, not just in a demo.”

What the scores actually show

The headline finding: most agents pass most scenarios with no extra guidance at all. Source In the Build stage — where an agent creates something from scratch — two models scored 100% unaided. Skills closed the remaining gap for others:

AI Model Build stage, no skills Build stage, with skills
Opus 5 100%
Kimi K3 100%
Sonnet 5 78% 100%
GPT-5.6 Sol 89% 100%
GPT-5.4 mini 78% 89%

As reported in the Supabase Evals announcement.

One more finding should give every business owner pause: documentation habits vary wildly. One model read roughly 8 documentation pages per task. Another checked the documentation in fewer than 40% of scenarios. Source In plain terms: some AI agents build things without reading how the tools are meant to be used. That’s precisely the behaviour that produces security holes.

Practical takeaways for your business

  • Ask your vendors how they test AI-generated work. If AI builds anything that touches customer data, there should be a verification process for security rules and database structure — reviewed by a human who knows what they’re looking at.
  • Check public leaderboards. For the first time, you can look at a public leaderboard and see how different AI models perform on real development tasks. You don’t need the technical details. You need to know if your vendor’s tools score well or don’t appear at all. Source
  • Never skip the human review step. The benchmark allows one retry before grading. Your business should, too. Anything AI builds should be reviewed by a person before it touches real customer data.
  • Start small. Let AI handle a low-risk internal task first — an internal dashboard, a reporting tool, a non-critical workflow. Watch how it’s built and reviewed before trusting it with anything sensitive.
  • Insist on documented guidance. The benchmark showed that agents given written guidance produce better results. A vendor who documents your systems for their AI will get better outcomes than one who doesn’t.

The bigger picture

Here’s what’s really happening. Benchmarks like Supabase Evals are the first step toward verifiable AI work. Right now, when a vendor says “AI built this,” you have no independent way to check whether that’s meaningful. Public benchmarks change that. They create a shared record of what AI agents can and cannot do, tested on real tasks in real environments.

For Malaysian SMEs, the opportunity is significant. AI that can reliably build and fix software means the barrier to getting custom tools built is falling. Things that once required a development team — custom dashboards, automated workflows, integrated systems — are becoming within reach. But the flip side is the responsibility we discussed: human oversight doesn’t disappear. It shifts.

The business owners who thrive in the next few years won’t be the ones who learn to code. They’ll be the ones who learn to direct, check, and verify AI work. They’ll ask sharper questions of their vendors, demand testing and review processes, and understand just enough about security to know when something is too risky. Benchmarks like this are the tool for asking those questions with confidence.

The AI that builds your next system might be genuinely good. Now, for the first time, you have

Ready to Streamline Your Operations?

Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →