AI Code Testing Boom: What It Means for Your Business

AI Code Testing Boom: What It Means for Your Business — featured image

by

Your Software Keeps Breaking — and AI Is About to Make It Happen Faster

Think about the last time one of your digital tools failed you. Maybe your online ordering form swallowed a payment. Maybe your inventory spreadsheet duplicated the same entry across 14 rows. Maybe your WhatsApp Business auto-reply sent the wrong price list to a long-time client.

Now imagine the companies building those tools are under growing pressure to ship new features quickly — and they’re using AI to write large portions of the code. That’s not a distant Silicon Valley story. It’s happening right now in the very software your business depends on.

A startup called Blacksmith, which helps companies test and verify software before it reaches users, just saw its valuation jump nearly 10x in less than a year. That’s not a vanity metric. It’s a warning signal: AI can write code much faster than anyone can check it, and the gap is growing.

Here’s why you should care, even if you’ll never type a line of code in your life.

TL;DR

  • AI coding tools let software teams produce code at record speed, but speed creates a testing bottleneck.
  • Blacksmith, which verifies AI-generated code, grew from 700 to more than 5,000 customers in under a year — proof that the problem is widespread.
  • For you: reliability of your digital tools matters more than how fast they’re built. Ask better questions of vendors, and validate your own AI outputs before they go live.

What This Means: More Code, More Gatekeeping

Let’s translate this into plain language. In the past, a human developer wrote code slowly — line by line, understanding every piece. Mistakes happened, but they were usually contained and caught before release.

Today, tools like Cursor, OpenAI’s Codex, and Anthropic’s Claude Code let developers generate hundreds of lines of code in minutes. The catch: AI code can be confidently wrong. It might pass a quick review but fail in real use — on an older phone, a slow connection, or an unusual browser.

Blacksmith, founded in 2024, started as a service that runs the building and testing of software for other companies, and has since expanded into an AI agent that fixes failed checks itself. Its latest Series B round was led by Peak XV Partners, with GV and Y Combinator also participating. Its founder summed up the industry’s pain in one sentence:

“Validating code is still a bottleneck, and it’s an even bigger bottleneck because people are writing even more.” — Aditya Jayaprakash, co-founder and CEO of Blacksmith

Think of it like your own business. If a machine let your team produce 10 times more product in a day, you would also need 10 times more quality checks — plus a system for catching defects before delivery to customers. Software is no different.

How This Applies to Malaysian SMEs

Your business runs on software whether you reflect on it or not. The e-commerce platform you sell through, the accounting system your bookkeeper uses, the booking tool your customers depend on, the e-invoicing system you’ll be filing through — all of it is built by developers using modern tools, increasingly with AI assistance.

Here’s the practical angle: when a vendor uses AI to ship features faster, you inherit the risk. A bug in their code becomes a lost order for you. Consider that Blacksmith’s customers include Mercury, Supabase, Clerk, Ashby, and Expensify — businesses with serious engineering teams, yet they still choose to use a dedicated testing platform as part of their workflow. So before you subscribe to a new tool or renew an existing one, ask directly: “How do you test your software?” Good vendors talk about automated testing, staging environments, and quality checkpoints. Vague answers should raise a red flag. The emergence of a thriving testing industry around AI code should tell you that quality is not guaranteed.

The second angle is personal. You probably use AI in your own daily work — drafting product descriptions, writing marketing posts, replying to customer enquiries, or summarising sales data. AI output looks polished but can contain invented figures, wrong dates, or off-brand wording. Your team now carries a new responsibility: reviewing and validating before anything goes public. The principle is the same one Blacksmith applies to code, applied to your everyday operations.

Third, consider how your workflows are designed. If AI lets your team produce 10 times more content or handle 10 times more enquiries, you have shifted the bottleneck to reviewing and approval. Malaysian SMEs often run lean — one person juggling multiple hats. That means you need a simple, repeatable validation process. For example, a standard checklist before any customer-facing post goes out, or a second-person review for anything involving numbers, prices, or legal commitments. The businesses that build this discipline early will benefit from AI without suffering its mistakes.

There’s also the regulatory layer. With e-invoicing now mandatory for more Malaysian businesses, the accuracy of your software is not just convenient — it is a compliance requirement. A small error in invoicing logic could mean incorrect submissions and unnecessary headaches. Software reliability has moved from an IT concern to a business risk, and it deserves the same attention you give to cash flow and customer service.

What the Numbers Tell Us

Metric Under a Year Ago Now Growth
Customers served 700+ 5,000+ ~7x
Team size 10 ~30 ~3x
Company valuation Series A round Nearly 10x higher 10x

All of this happened in less than 12 months, while serving the companies listed above. If those teams need help validating code, imagine how exposed a small business is when relying on off-the-shelf tools built by small teams racing to ship features.

Practical Takeaways: A Simple Checklist

  • Question new tools before adopting them. Run a pilot for two to four weeks. Test with real scenarios: unusual addresses, older phones, different browsers.
  • Ask vendors how they test. If they cannot explain their quality process, treat that as a warning sign.
  • Create a review checklist for AI-generated content. Names, numbers, prices, dates — always verify before publishing.
  • Never auto-schedule AI output without review. One wrong price sent to 500 customers is hard to take back.
  • For compliance-related work (invoicing, taxes, contracts), insist on a human check no matter how polished the AI output looks.

The Bigger Picture

This trend signals a broader shift: the software industry is moving from “how fast can we build it” to “how sure are we that it works.” For Malaysian SMEs, that is encouraging news. It means more reliable tools should reach the market, and the specialists checking AI code will keep improving the quality of everyday platforms.

But it also means you must become a better selector — and a better validator of your own outputs. The companies that succeed in the next five years will not be the ones using the most AI. They will be the ones with the best systems for checking their own work, whether the work came from a human or a machine.

The good news: you can build that discipline today, without writing a single line of code.

Ready to Streamline Your Operations?

Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →