Self-Improving AI: What Malaysian SMEs Should Do Now

Self-Improving AI: What Malaysian SMEs Should Do Now — featured image

by

AI That Improves Its Own Work Is Moving Closer

You may already use AI to draft customer replies, summarise documents, prepare marketing ideas or organise internal information. The difficult part is not getting an answer. It is knowing whether the answer is reliable, safe and suitable for your business.

That problem becomes more important as AI systems begin helping to improve other AI systems. A new Anthropic research paper describes an automated system that searched research literature, proposed training methods, tested them and kept the methods that improved results. The work offers an early view of AI that can assist with its own improvement, while also showing why human checks remain essential.

TL;DR: AI systems are beginning to test and improve model behaviour automatically. For your business, the practical lesson is to build clear rules, test AI against real tasks and keep people responsible for important decisions.

The research is not proof that AI can independently replace every researcher or business function. It is better understood as a warning and an opportunity: the businesses that develop disciplined AI processes will be in a stronger position than those that simply ask a chatbot to “do everything”.

What This Means

Anthropic’s paper describes an “Automated Alignment Researcher” that worked on benchmarks involving specific misaligned behaviours. In plain language, the system was given tests designed to measure whether an AI model followed intended safety and behaviour requirements. It searched available literature, suggested an approach, trained the model for a limited period and repeated the process.

The reported system improved performance across all 10 alignment benchmarks without reducing overall performance, according to the source article. TechCrunch reported the results and described the automated research process. The system also retained methods that worked and discarded methods that did not, allowing it to run many experiments in sequence.

This is sometimes described as a step towards recursive self-improvement. The phrase sounds dramatic, but the practical meaning is straightforward: one AI system helps find better ways to train, evaluate or improve another AI system. The improvement cycle may eventually cover more than safety training, including reasoning, coding, research and task performance.

However, the system is only as good as the tests and goals it receives. If your benchmark measures the wrong thing, an AI may become better at passing the test without becoming better at serving customers. This is similar to a sales team that learns to increase the number of calls but ignores whether customers receive useful answers.

Key insight: AI does not automatically know what “better” means. You must define the outcome, create useful tests and review whether the results match your real business goals.

What the Research Shows

The source article highlights several features that business owners should understand. The automated system copied parts of a traditional research workflow: reviewing available knowledge, proposing an experiment, testing the proposal and refining the next attempt. The paper states that the best automated method outperformed human-proposed methods on average within six hours. The reported comparison is described in the TechCrunch source.

It also reports that human researchers still have an important role in creating and maintaining the benchmarks, selecting suitable research material and checking whether the system’s results reflect genuine improvement. This limitation matters for SMEs because business data is often incomplete, inconsistent or spread across messaging apps, spreadsheets and accounting systems.

Research detail Business lesson
10 alignment benchmarks were tested Define several practical checks instead of relying on one general AI score
Training runs lasted 30 minutes per iteration Use short, repeatable tests before making a workflow permanent
Methods were retained or discarded based on results Keep a record of what works and remove weak prompts or processes
Benchmark quality remained a limitation Review whether your tests reflect actual customer and operational needs

These figures and limitations come from the source article’s summary of Anthropic’s paper. Read the full source report for the research context.

How This Applies to Malaysian SMEs

Suppose you run a service company and use AI to draft WhatsApp replies to enquiries. Instead of checking only whether the message sounds polite, create a small evaluation set: Does it answer the customer’s question? Does it avoid promising unavailable services? Does it use the correct business hours? Does it escalate complaints to a human? You can test the same 20 or 30 common enquiries whenever you change your prompt, knowledge base or automation.

For a retailer, AI may help classify customer questions, recommend products or prepare follow-up messages. Your benchmark could check product availability, delivery areas, return rules and language quality. Malaysian customers may move between Bahasa Malaysia and English, while some industries also handle Mandarin or Tamil communication. A useful test should reflect the actual language mix and common questions your team receives, not an idealised example written by a software vendor.

For a small accounting, construction or professional services firm, AI may summarise documents and prepare internal task lists. Here, the most important checks concern accuracy and confidentiality. Does the summary include the correct dates? Are missing documents clearly identified? Does it avoid inventing figures? Is sensitive client information kept within an approved system? An automated workflow should not be allowed to send a final tax, legal or contractual conclusion without human review.

Restaurants, clinics, tuition centres and repair businesses can also apply the same idea. You could test whether an AI assistant records bookings correctly, recognises urgent requests, asks for missing details and routes unusual cases to a staff member. The goal is not to make the AI appear clever. The goal is to reduce repetitive work without creating a new stream of errors.

For Malaysian SMEs, the most realistic first step is not building a self-improving model. It is creating a self-improving process. Each week, collect examples of good and bad AI outputs, update your instructions, add new tests and review the results. This gives your team a controlled improvement loop without requiring technical expertise.

Practical Takeaways

  • Start with one workflow: Choose a repetitive task such as enquiry replies, document summaries or appointment handling.
  • Write down the desired result: Describe what a correct answer must contain and what it must never do.
  • Create a test set: Keep real examples, with private information removed, covering normal, difficult and unusual cases.
  • Use several checks: Measure accuracy, completeness, tone, escalation behaviour and compliance with your internal rules.
  • Keep humans in charge: Require approval for legal, financial, medical, employment or customer-complaint decisions.
  • Record changes: Note which prompt, data source or workflow version produced each result.
  • Review failure patterns: Do not only celebrate successful outputs; examine where the system misunderstood customers.
  • Control access: Give AI tools only the information and permissions they need for the assigned task.
  • Train your staff: Teach employees when to trust an AI suggestion, when to verify it and when to escalate.

A Simple 30-Day Starting Plan

  1. Days 1–7: Select one process and collect 20 representative examples. Remove confidential information before using them for testing.
  2. Days 8–14: Define a pass-and-fail checklist. Ask two staff members to review the examples so the standard is not based on one person’s opinion.
  3. Days 15–21: Run the AI workflow, record errors and improve the instructions. Separate harmless style issues from serious accuracy or privacy issues.
  4. Days 22–30: Repeat the test, compare results and decide whether the workflow is ready for limited staff use. Set a review date rather than treating it as finished forever.

The Bigger Picture

Self-improving AI may eventually make the technology cycle faster: models could generate experiments, evaluate outcomes and suggest better training methods with less direct intervention. For business owners, this means AI tools may change more frequently and improve in ways that are difficult to see from a product demo.

That creates a management responsibility. You should know which AI tools your team uses, what information they receive, what decisions they influence and how their performance is checked. A tool that improves its responses may still become unsuitable if your products, policies or customer expectations change.

The long-term advantage will not come from using the most impressive AI feature. It will come from having reliable business data, clear operating rules and a habit of reviewing results. When an AI system improves, your business should be able to tell what improved, why it improved and whether the change is safe to adopt.

For now, treat self-improving AI as a signal to strengthen your foundations. Start with a narrow workflow, define success in practical terms and keep accountable people involved. That approach will help you benefit from better AI tools without handing over decisions your business cannot afford to get wrong.

Ready to Streamline Your Operations?

Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →