When a Customer Needs You to Get the Details Right
If your business handles customer calls, you already know that a pleasant voice is not enough. A caller may give an order number, email address, delivery reference, phone number, appointment date, or support ticket code. If your automated agent reads even one digit incorrectly, the conversation can create more work instead of removing it.
Speed matters too. Long pauses make callers wonder whether the line has dropped. They may repeat themselves, hang up, or ask to speak to a human. For a small team, those failed interactions can interrupt staff and make routine enquiries harder to manage.
Gradium AI has released a text-to-speech model designed to improve both accuracy on difficult information and response speed. The model became the default across its API and Studio on August 31, 2026, while existing voices and custom clones continue working without migration, according to MarkTechPost.
TL;DR
Gradium reports an 81.0% human-rated pass rate on a 500-sentence test covering numbers, emails, dates, acronyms, and realistic service scenarios, as reported by MarkTechPost.
It also reports 216 milliseconds at the median time to first audio, which may help voice agents respond with fewer awkward pauses. You should still test it with your own Malaysian customer conversations before putting it in front of callers.
What This Means
Text-to-speech, or TTS, is the technology that turns written AI responses into spoken audio. A basic TTS system may sound natural while still struggling with information that is important to a business. Numbers, email addresses, product codes, dates, abbreviations, and mixed letters can be difficult to pronounce consistently.
Gradium says its new model was evaluated using 500 sentences across five languages: English, German, French, Spanish, and Portuguese. The evaluation included ten criteria. Seven focused on individual details such as spelling, acronyms, alphanumeric tokens, dates, regular numbers, large and floating numbers, and email addresses. Three combined these details into scenarios involving orders, IT tickets, and claims, according to MarkTechPost.
The reported 81.0% pass rate means a sentence passed only when a human rater judged every required element to be pronounced correctly and completely. One missing or misread digit caused the sentence to fail. Gradium compared its result with Cartesia Sonic 3.6 at 75.1%, ElevenLabs v3 Conversational at 65.4%, Fish Audio S2.1 Pro at 49.5%, and Inworld TTS 1.5 Max at 46.5%, based on the same report.
For response speed, Gradium reports a median time to first audio of 216 milliseconds across Coval benchmark testing. It also reports a 30-millisecond interquartile spread across 480 runs, while Cartesia Sonic 3.6 recorded a 454-millisecond median and a 165-millisecond spread in the comparison cited by MarkTechPost.
The useful question is not whether an AI voice sounds human. It is whether customers can trust what it says when the information contains digits, codes, names, and dates.
How This Applies to Malaysian SMEs
1. Delivery and order enquiries
If you sell food, skincare, clothing, spare parts, or other products, customers often call to ask about order status. An AI voice agent could read an order reference, confirm a delivery date, or repeat a phone number for a callback. This is where pronunciation accuracy matters more than an impressive demonstration voice.
Before deploying such a system, prepare a test list based on your actual orders. Include short and long reference codes, repeated digits, hyphens, letters that sound similar, and dates written in the format your staff use. Ask several people to listen without seeing the written text. If they cannot reliably write down the same information, the workflow needs adjustment.
2. Clinics, salons, workshops, and appointment-based services
A small clinic or salon may receive calls asking about appointment times, branch locations, registration details, or rescheduling. A workshop may need to confirm vehicle registration numbers, service dates, and job references. An automated voice agent can handle basic confirmation while your staff focus on customers already at the premises.
For Malaysian customers, you may also need to decide when the agent should use English, Bahasa Malaysia, Mandarin, Tamil, or a human handover. Gradium’s published hard-case evaluation covers English, German, French, Spanish, and Portuguese, but not Bahasa Malaysia, Mandarin, or Tamil, according to MarkTechPost. That does not prove it will perform poorly in local languages; it means you should test those languages yourself rather than assume the published result applies.
3. Property, insurance, and service enquiries
Property agents, insurance advisers, repair companies, and maintenance providers often deal with addresses, policy numbers, claim references, appointment windows, and identification details. These are useful situations for a voice agent, but they also carry a higher risk when information is misheard.
You can begin with low-risk tasks: collect the caller’s preferred contact time, answer common questions, provide office hours, and create a request for staff follow-up. Keep sensitive or complex decisions with a trained employee. The AI should repeat important details slowly, ask the caller to confirm them, and send the same information through an approved written channel where appropriate.
4. B2B enquiries and internal coordination
If you supply products to other businesses, callers may mention stock-keeping units, purchase order numbers, invoice references, or technical model names. A voice agent can capture the enquiry outside office hours and route it to the right person. Faster first audio may make the interaction feel more responsive, but accuracy remains the priority.
Use structured fields behind the conversation. Instead of asking the AI to remember a long paragraph, design the workflow to capture the customer name, company, reference number, request type, and follow-up deadline separately. This gives your staff a clearer record to review and reduces the chance that a useful detail is buried in a transcript.
A Quick Comparison of the Reported Results
| Model | Hard-case pass rate | Reported timing detail |
|---|---|---|
| Gradium TTS | 81.0% | 216 ms median time to first audio |
| Cartesia Sonic 3.6 | 75.1% | 454 ms median time to first audio |
| ElevenLabs v3 Conversational | 65.4% | 329 ms median time to first audio |
| Fish Audio S2.1 Pro | 49.5% | 291 ms median time to first audio |
| Inworld TTS 1.5 Max | 46.5% | 166 ms median time to first audio |
These figures come from the vendor-reported comparison described by MarkTechPost. They are useful for understanding the claim, but they are not a guarantee for your call setup, network conditions, accents, language mix, or customer vocabulary.
Practical Takeaways
- Start with one narrow call type. Choose order status, appointment confirmation, or basic enquiry capture rather than trying to automate every conversation.
- Build a hard-case test set. Include at least 20 examples containing phone numbers, emails, dates, codes, decimals, abbreviations, and names used by your customers.
- Test Malaysian speech patterns. Check English, Bahasa Malaysia, code-switching, local names, and common pronunciation differences.
- Make the agent repeat critical details. Ask the caller to confirm the information instead of treating the first reading as final.
- Provide a human escape route. The caller should be able to request staff help without repeating the entire conversation.
- Review transcripts regularly. Track which information is most often corrected, misunderstood, or abandoned.
- Protect customer information. Decide what the system may read aloud and what should be handled only after proper verification.
- Measure completed tasks. Look at successful bookings, correctly captured references, and resolved enquiries, not just voice quality.
The Bigger Picture
Voice automation is becoming more practical for small businesses because the quality target is changing. You do not need an agent that handles every unusual conversation. You need one that performs a limited set of repetitive tasks accurately, responds promptly, and knows when to involve a person.
That makes evaluation more important than marketing demonstrations. Open datasets, such as the 500-sentence evaluation set Gradium says it released on Hugging Face under a CC BY 4.0 licence, can help teams understand how a model is tested, as reported by MarkTechPost. You should still create a second test set from your own business because your product names, customer languages, codes, and call patterns will be different.
For a Malaysian SME, the sensible path is gradual. Keep your existing customer service process, add automation around predictable questions, and monitor every important detail that the system reads aloud. If the agent saves staff time without creating correction work, expand it carefully. If it causes confusion, narrow its role and improve the handover process.
The strongest benefit of faster, more accurate TTS is not that customers are impressed by an AI voice. It is that your team can respond consistently to routine calls while reserving human attention for cases that genuinely need judgement, empathy, or local knowledge.
Ready to Streamline Your Operations?
Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →
