Why LLM Scores Can Mislead Your SME’s AI Decisions

Why LLM Scores Can Mislead Your SME’s AI Decisions — featured image

by

Your AI test score may not tell the whole story

You may be comparing AI tools for customer support, document drafting, sales follow-up, internal knowledge search, or workflow automation. The usual approach is simple: look at a benchmark score, compare several models, and choose the one with the highest number.

That approach feels practical, but the number may hide what the model is actually good at. A test advertised as measuring safety, reasoning, or instruction-following may contain questions that also depend on reading ability, background knowledge, context tracking, or the ability to interpret tricky wording.

This matters when you are choosing an AI system for your business. A model that performs well on a general test may still struggle with mixed-language customer messages, product-specific instructions, sensitive documents, or the exact approval rules your staff follow.

TL;DR: BenchMIRT shows that one LLM benchmark score can combine several different abilities. You should assess AI using tasks that resemble your real Malaysian SME workflows, not a single headline score.

Research from the Allen Institute for AI analysed 100 large language models across 16 benchmarks and more than 34,000 questions using BenchMIRT. The method identified two major underlying dimensions: safety and general reasoning. Source: Hugging Face and Allen Institute for AI

What This Means

Think of an AI benchmark as an exam. A benchmark may claim to test one subject, but individual questions can require several skills at once. A question intended to measure bias, for example, may also require the model to understand who is speaking, track relationships, and reason from evidence rather than assumptions.

BenchMIRT examines performance at the level of individual questions. It uses a method called multidimensional Item Response Theory, or MIRT. This approach comes from psychometrics, the field that studies how tests measure ability. It considers both the capability of the model and the characteristics of each question.

In plain language, BenchMIRT asks:

  • How difficult is this question?
  • Which capability does it actually test?
  • Does it separate stronger models from weaker ones?
  • Is the overall benchmark score being influenced by another skill?

The study found that several benchmarks broadly matched their intended purpose. Reasoning tests tended to reflect reasoning ability, while harmful-request and jailbreak tests tended to reflect safety behaviour. However, some results were more complicated.

For example, BBQ is commonly used to evaluate social bias, yet its results aligned more strongly with general reasoning in the analysis. A low score could therefore reflect difficulty understanding or reasoning through the question, not only a problem with bias-related behaviour.

WMDP, a benchmark covering dangerous dual-use knowledge in areas such as biology, chemistry, and cybersecurity, also showed an unexpected pattern. Stronger general reasoning was associated with lower WMDP scores because the desired behaviour was to refuse or fail to provide dangerous information. A higher score did not automatically mean a better general-purpose model.

HarmBench showed another issue: different groups of questions within one benchmark could measure different signals. Standard harmful requests and contextual harmful requests were more closely associated with safety, while copyright-related questions aligned more closely with general reasoning. Source: BenchMIRT findings

A benchmark score is evidence, not a complete description of what an AI model can do.

Why This Matters When You Choose AI

As a business owner, you are not buying an abstract reasoning score. You are choosing whether an AI tool can draft a quotation accurately, classify an incoming enquiry, summarise a customer complaint, follow your approval policy, or avoid exposing confidential information.

These tasks involve several abilities at once. The model must understand the request, identify relevant details, follow instructions, use your business context, handle ambiguity, and avoid an unsafe or unauthorised response. A general benchmark may test only part of this combination.

BenchMIRT also found that keeping only 10% of questions could generally preserve nearly the same picture of which models were stronger or weaker on the underlying capabilities, while retaining 50% often matched the full benchmark even more closely. Its method predicted whether a model would answer an unseen question correctly 79% of the time, compared with 70% for a simpler overall-score approach. Source: BenchMIRT experiments

This does not mean you should ignore benchmarks. It means you should read them carefully and add your own evaluation. A shorter set of highly relevant tests may be more useful than a large collection of unrelated questions.

How This Applies to Malaysian SMEs

1. Customer service and WhatsApp enquiries

Suppose your business receives enquiries in Bahasa Malaysia, English, Mandarin, Tamil, or informal mixed-language messages. A model may score well on general reasoning but misunderstand local phrasing, abbreviations, or customer intent. It may also answer a simple product question correctly while failing when the customer provides several details in one message.

You should test realistic examples from your own enquiry history, with confidential information removed. Include short messages, unclear messages, follow-up questions, and customers who switch languages. Measure whether the AI identifies the correct product, asks for missing information, uses the right tone, and escalates sensitive cases to a staff member.

2. Quotations, invoices, and order processing

An AI assistant handling sales documents must do more than extract text. It may need to distinguish delivery instructions from billing details, identify quantities, spot missing information, and follow your internal approval rules. A model can appear capable in a general document test but still confuse item codes, units, dates, or customer names.

Create a test set using common quotation and purchase-order formats from your operations. Include scanned documents, tables, handwritten notes where relevant, and documents with incomplete fields. Check every output against a known answer. For a Malaysian business, you may also need to test local address formats, registration details, tax-related fields, and references to bank or payment instructions without allowing the AI to approve transactions on its own.

3. Internal knowledge and staff support

If you use AI to answer questions about leave procedures, product specifications, service policies, or standard operating procedures, a high reasoning score is not enough. The system must retrieve the correct internal source and avoid inventing an answer when the information is missing.

Test questions that require the AI to compare two procedures, locate a small detail, and say “I do not know” when your documents do not provide an answer. Include outdated documents to see whether the system follows the latest approved version. This is especially important when several branches or departments use different procedures.

4. Marketing content and sales follow-up

An AI tool may write fluent copy but still make unsupported claims, misunderstand your target customer, or use a tone that does not suit your brand. Benchmark performance on general language tasks cannot tell you whether a follow-up message is appropriate for a local customer or whether a product description accurately reflects what you sell.

Give the model approved product facts, prohibited claims, preferred languages, and examples of acceptable messages. Then evaluate factual accuracy, clarity, tone, and whether the suggested next step is suitable. You can ask several staff members to rate the outputs using the same checklist.

A practical evaluation method for your business

Evaluation area Example business test What to check
Instruction following Draft a reply using a fixed format and approval rule Does it follow every required step?
Reasoning Choose the correct product or service from several conditions Does it use all relevant details?
Safety Handle a request for confidential or unauthorised information Does it refuse or escalate appropriately?
Local language handling Respond to mixed Bahasa Malaysia and English messages Does it preserve meaning and tone?
Business accuracy Summarise an order, complaint, or quotation Does it avoid missing or invented details?

The table above is a practical starting structure rather than a universal scoring system. You should adapt the examples to your industry and record the results consistently.

Practical Takeaways

  • Do not choose an AI model from one benchmark score alone.
  • Ask what capabilities the benchmark was designed to measure.
  • Check whether the benchmark mixes safety, reasoning, language, and knowledge tasks.
  • Build a small evaluation set from real business scenarios with sensitive details removed.
  • Include Malaysian language patterns, local document formats, and your normal customer questions.
  • Test both successful answers and refusal or escalation behaviour.
  • Separate accuracy, instruction-following, safety, and tone instead of using one combined rating.
  • Review difficult examples individually; averages can hide important failures.
  • Retest the AI when your documents, products, policies, or workflow change.
  • Keep a human approval step for high-impact actions such as payments, legal commitments, hiring decisions, and sensitive customer communication.

The Bigger Picture

BenchMIRT points towards a more careful way of evaluating AI. Instead of treating a benchmark as a single score, researchers can examine which questions provide useful information and which abilities are influencing the result. That should lead to evaluations that are smaller, clearer, and more relevant to the decision being made.

For SMEs, the long-term lesson is straightforward: your own workflow data and test cases may be more useful than a generic leaderboard. A public benchmark helps you understand broad model capabilities, but your business has its own vocabulary, documents, customers, risks, and approval rules.

You do not need a research team to apply this principle. Start with 30 to 50 representative tasks, label what a correct answer looks like, and compare tools using the same examples. Separate “got the facts right” from “followed the instruction” and “handled the risk correctly.” This gives you a much clearer basis for deciding where AI can assist your team.

The best question is not “Which model has the highest score?” It is “Which model performs reliably on the tasks that matter to my business, and what does it do when it is uncertain?” That shift can help you make calmer, safer, and more practical automation decisions.

Ready to Streamline Your Operations?

Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →