Why AI Benchmark Scores May Mislead Your SME

Why AI Benchmark Scores May Mislead Your SME — featured image

by

A Better Way to Judge AI Tools for Your Business

If you are choosing an AI tool for customer service, sales, document processing or internal operations, you may be tempted to compare models using one headline score. A higher score appears to mean a better tool. But new research suggests that many large language model (LLM) benchmarks measure several abilities at once, making a single number difficult to interpret.

This matters to Malaysian SME owners because your business does not need “the best AI” in the abstract. You need an AI system that can understand your products, follow instructions, handle Bahasa Malaysia or English, protect sensitive information and produce dependable results for your actual workflow. A model that performs well on a public reasoning test may still be unsuitable for replying to WhatsApp enquiries, summarising invoices or preparing a customer follow-up.

What Happened

The Allen Institute for AI introduced BenchMIRT, a method for auditing LLM benchmarks at the level of individual questions and tasks. The approach examines what each prompt is really testing instead of relying only on the benchmark’s overall score. The source article was published on 1 September 2026 by Hugging Face and Ai2. Read the original article.

BenchMIRT builds on Item Response Theory, a method originally used to study test performance. It uses a multidimensional approach, known as MIRT, to estimate both the capability of a model and the difficulty of individual questions. In the study, the researchers analysed results from 100 LLMs across 16 benchmarks and more than 34,000 questions. See the BenchMIRT methodology.

Without being told what each benchmark was intended to measure, BenchMIRT repeatedly identified two major dimensions: safety and general reasoning. This suggests that benchmark results can contain overlapping signals. A model may score poorly not because it lacks safe behaviour, but because it struggles to understand the wording, context or reasoning required by a particular question. Read the technical report.

Why One Score Can Hide Important Details

The research found that some benchmarks broadly matched their intended purpose, while others were more complicated. BBQ, a benchmark designed to assess social bias, was found to align more closely with general reasoning than safety in the analysis. A low score could therefore reflect difficulty following the scenario or identifying the relevant evidence, not only a problematic safety response.

WMDP, which tests dangerous dual-use knowledge in areas such as biology, chemistry and cybersecurity, also showed an unexpected relationship. Its preferred behaviour is for a model to refuse or avoid providing harmful information. The analysis found that stronger reasoning was associated with lower WMDP scores because the benchmark treats refusal or failure to provide dangerous information as the desired result. Review the benchmark findings.

HarmBench showed another issue: different groups of questions within one benchmark can measure different signals. Prompts involving harmful requests were more closely related to safety, while some copyright-related questions were more closely linked to general reasoning. Combining them into one average score can make the result less clear.

A benchmark score is not a complete description of an AI model. It is a measurement produced by a particular set of questions, and those questions may test more than the benchmark label suggests.

Why This Matters for Malaysian SMEs

For a Malaysian SME, the practical lesson is simple: evaluate AI against your real work, not only a vendor’s ranking table. Suppose you operate a retail business in Shah Alam and want an AI assistant to answer customer questions. The important tests may include whether it understands “boleh COD ke?”, distinguishes delivery areas, handles mixed Bahasa Malaysia and English, and follows your refund policy accurately. A general reasoning score cannot tell you all of that.

The same applies to a construction company, accounting practice, clinic, distributor or online seller. An AI tool for quotation preparation should extract product quantities correctly, identify missing information and ask for clarification. A tool for staff onboarding should follow your internal rules. A tool that reads invoices should recognise Malaysian formats, supplier names and tax-related fields. These are workflow capabilities, and they should be measured directly.

You should also separate safety from usefulness. An AI system that refuses too many harmless requests may look safe but create frustration for your staff. For example, a customer service assistant should reject requests for confidential information, but it should still answer normal questions about opening hours, delivery status and product specifications. If you combine refusal behaviour and general task performance into one score, you may miss this balance.

Language and context deserve their own checks. Malaysia’s business communication commonly moves between English, Bahasa Malaysia, Mandarin or Tamil depending on the audience and industry. A model may perform strongly on an English benchmark but misunderstand local abbreviations, informal messages or code-switching. Build test prompts from anonymised conversations and documents that resemble your daily operations.

A Practical AI Evaluation Checklist

Area to test Example for your business What to record
Instruction following Prepare a reply using your preferred tone and format Whether every requirement is followed
Local language handling Answer a mixed Bahasa Malaysia-English enquiry Accuracy, clarity and suitable terminology
Reasoning Compare delivery options or identify a missing quotation detail Whether the conclusion is correct and explainable
Safety Handle a request for confidential customer or staff information Whether it refuses appropriately without blocking normal work
Reliability Repeat similar tasks using different wording Consistency of the answers
Human handover Escalate a complaint or uncertain request Whether the issue reaches the right person

Use a small, representative test set before introducing an AI tool to customers or staff. Include easy, medium and difficult cases. Add unusual cases, such as incomplete information, spelling mistakes and messages containing several requests. Ask at least two people who understand the workflow to review the results. Record not only whether the answer is correct, but also whether it is useful, safe and easy to verify.

The Bigger Picture

BenchMIRT points towards a more careful approach to AI evaluation. The researchers reported that retaining around 10% of selected questions generally preserved a similar view of which models were stronger or weaker on the underlying capabilities. They also reported that the method predicted held-out question performance correctly 79% of the time, compared with 70% for a simpler overall-score approach. See the reported results.

For SMEs, this does not mean you need to build a research-grade benchmark. It means you can learn from the principle: fewer, better-designed tests may be more useful than a large collection of generic questions. Select the tasks that matter most to your operation, make the expected answer clear and review failures by category.

You can classify failures into several groups: knowledge gaps, poor reasoning, unclear instructions, language misunderstanding, unsafe disclosure or inability to recognise when human help is needed. This tells you what to improve. You may need better internal documents, clearer prompts, stronger access controls, staff training or a different model for a particular task.

AI adoption should therefore be treated as an operational project rather than a one-time software purchase. Start with one workflow, define success, test real examples and monitor performance after launch. Keep a human approval step for sensitive actions such as refunds, employment decisions, legal communications and changes to customer records.

The headline benchmark still has value as an initial comparison. However, BenchMIRT’s central message is useful for every business owner: ask what a score is actually measuring. For your SME, the winning AI tool is not necessarily the one with the highest public ranking. It is the one that performs reliably on your customers, your documents, your languages and your daily decisions.

What You Can Do Next

  • List one repetitive workflow where AI could assist your team.
  • Collect representative, anonymised examples from that workflow.
  • Separate tests for reasoning, language, safety and instruction following.
  • Compare several tools using the same questions and scoring method.
  • Review failures with the staff who handle the process every day.
  • Launch gradually with human oversight and a clear escalation process.

Benchmark scores can start a conversation, but your own evidence should guide the final decision. By testing AI at the level of the actual task, you can reduce surprises and choose automation that supports the way your Malaysian business really operates.

Ready to Streamline Your Operations?

Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →