Why a Live AI Search Benchmark Matters to Your Business
If you use AI to research suppliers, monitor competitors, answer customer questions or prepare business reports, search quality matters more than the chatbot’s writing style. An AI system can produce a polished answer while missing the newest announcement, relying on weak evidence or failing to find a rare but important fact.
That is why a new open-source project called NEEDLE deserves attention from Malaysian SME owners. Developed by Keenable AI, NEEDLE tests web search APIs using query sets that are rebuilt regularly from fresh public information, instead of relying on a fixed test dataset. The approach is designed to show whether a search system can find current and difficult information in realistic conditions. Source
You do not need to install the benchmark to benefit from the lesson. The practical takeaway is simple: before you trust an AI-powered workflow, test whether it retrieves the right information for your specific business, not merely whether it gives fluent responses.
What Happened
Keenable AI has open-sourced NEEDLE, which stands for News, Everyday, Expert, Deep-tail and Legal Evaluation. It is a live benchmark for comparing web search APIs used by AI agents. Its query sets are regenerated hourly for news and daily for finance, academic, rare-entity and legal searches. This reduces the risk that an AI system memorises a public answer key or becomes optimised for a permanently frozen collection of questions. Source
The benchmark draws from several public sources, including RSS feeds and Google Trends for news, SEC XBRL and business registries for finance, arXiv and Europe PMC for scholarly material, public agent-trajectory releases for rare queries, and CourtListener and eCFR material for legal searches. It then sends the same query to 15 search APIs under one protocol. Pages are not fetched or re-ranked, and each service is judged using its own titles, rankings and snippets, with evidence limited to 2,000 characters. Source
NEEDLE also creates an “ultimate” pooled oracle by combining the results found by all tested engines. This gives researchers an estimate of what the market collectively managed to retrieve. A wide gap between an individual engine and this ceiling suggests a ranking or retrieval problem. A low ceiling suggests that the available search systems failed to find strong evidence in the first place. Source
For your business, an AI answer is only as dependable as the search process behind it. A confident response does not prove that the system found the best source.
Why This Matters for Malaysian SMEs
Many Malaysian SMEs now use AI for tasks that depend on current information. A distributor may ask an assistant to identify new import requirements. A restaurant group may track food safety notices or changes affecting delivery operations. A software reseller may compare competitors’ product updates. A professional services firm may prepare a client briefing based on regulations, tenders or industry developments.
These tasks are not equally difficult. A question about a well-known company may be answered from common web pages. A question about a niche Malaysian supplier, a newly published agency notice or a technical product specification may require several searches and careful source checking. NEEDLE’s focus on fresh news, specialist information and “deep-tail” rare-entity queries reflects the type of research where a weak search system can create operational mistakes. Source
Consider a local sales team researching a prospect. If the AI finds only broad directory listings, it may miss a recent expansion, a new decision-maker or a procurement announcement. A hiring team could also overlook a relevant technical candidate if the search engine handles unusual names or specialist terms poorly. For you, the important question is not which provider has the highest global score. It is whether the system consistently finds the sources your staff need in Malaysia and in your industry.
| NEEDLE feature | Practical SME lesson |
|---|---|
| Fresh queries | Test AI workflows regularly because yesterday’s result may not reflect today’s information. |
| Multiple search categories | Evaluate news monitoring, supplier research, technical research and compliance separately. |
| Common query protocol | Compare tools using identical questions and the same evidence requirements. |
| Pooled “ultimate” ceiling | Distinguish a weak search provider from a problem that no available provider solves well. |
| Latency reporting | Measure response time when staff or agents perform many searches in one task. |
What the Published Results Show
The results reported for the seven-day period ending 28 August 2026 show that search performance varies by task. In finance, Exa recorded 0.910, Keenable 0.872, Perplexity 0.871 and Google 0.847, compared with an ultimate score of 0.965. Scholar searches showed a wider spread, from Keenable at 0.774 to Tavily at 0.310, against a ceiling of 0.869. Source
Deep-tail searches were harder. Exa reached 0.557 of the ultimate score, Keenable reached 0.470 and Bing reached 0.199. This is particularly relevant to SMEs because niche business research often includes uncommon product names, local firms, technical terms and incomplete descriptions. A system that performs well on general questions may struggle when your staff do not know the exact wording used in the source. Source
Speed also differed substantially. For the same reporting window, Keenable-realtime recorded 193 milliseconds at p50 and 284 milliseconds at p95, while Exa recorded 1,876 and 2,955 milliseconds. Bing recorded 2,767 and 9,381 milliseconds for those measurements. These figures come from the benchmark’s published results and should not be treated as a guarantee for your own network, plan or application setup. Source
How You Can Apply the Lesson
Start by creating a small evaluation set for your company. Write 20 to 30 real questions your staff ask repeatedly, such as “Which Malaysian suppliers provide this component?”, “What changed in the latest announcement?” or “Find the original source for this technical claim.” Include easy, current and deliberately obscure questions.
For every test, record the source URL, date, answer accuracy, missing information, duplicate results and response time. Ask the AI to show supporting links and flag uncertainty. Run the same test across the search tools you are considering. Do not judge only the final prose; inspect whether the cited page actually supports the statement.
- Use separate test sets for sales, operations, customer service and compliance.
- Include Malaysian company names, locations, product terms and Bahasa Malaysia queries where relevant.
- Require human approval for legal, safety, employment and regulatory decisions.
- Refresh your test questions when products, suppliers or regulations change.
- Keep a record of incorrect answers so your automation team can improve prompts and workflows.
The Bigger Picture
NEEDLE points to a broader change in AI evaluation. Static leaderboards are easy to understand, but they may not reflect the conditions under which your business operates. Information changes, queries are phrased differently and search agents often need to investigate several related questions before producing a useful answer.
For Malaysian SME owners, this means AI adoption should be managed like any other operational system. Define the task, measure the result, check the source and review performance over time. A live benchmark such as NEEDLE makes the industry’s limitations more visible, but you can apply the same discipline with a simple spreadsheet and a set of real business questions.
The best search tool for you may not be the one leading a general leaderboard. It is the one that finds dependable Malaysian and industry-specific information, provides traceable evidence, responds within your workflow’s time limit and fails safely when the answer cannot be verified.
Ready to Streamline Your Operations?
Technology moves fast. Your operations should keep up. AutoRunBiz builds AI systems that run your daily workflows — from WhatsApp order capture to accounting. Book a free 15-min ops audit →
