Smarter Document Extraction for Faster SME Workflows

Smarter Document Extraction for Faster SME Workflows — featured image

by

Why Document Extraction Matters to Your Business

As a Malaysian SME owner, you probably deal with documents every day: supplier invoices, purchase orders, delivery orders, signed agreements, staff forms, bank statements and customer applications. The problem is rarely finding a document. The real problem is getting useful information out of it without asking someone to read, check and retype everything manually.

A scanned invoice may contain the supplier name, invoice number, SST details, line items and payment terms, but those details are not automatically useful to your accounting or operations system. If your team copies the information into spreadsheets or accounting software, small errors can lead to payment delays, duplicate entries or unnecessary follow-up calls.

LandingAI’s new Agentic Document Extraction Gen2, announced on 9 September 2026, shows where document automation is heading. Instead of treating a document as a flat collection of text chunks, the system understands pages, sections, tables, figures, signatures and individual words. It also links extracted information back to the exact location where it appeared on the page.

TL;DR

LandingAI’s ADE Gen2 uses DPT-3 Pro and DPT-3 Verity to extract information from digital and scanned documents with more structure and traceability.

For your business, the important lesson is not the product name. It is the move towards document automation that can show what was extracted, where it came from and how confident the system is.

What This Means

Older document extraction tools often converted a file into a long stream of text. That can work for simple documents, but it becomes unreliable when layout carries meaning. In an invoice, for example, the supplier address, billing address, item description and total amount may all appear near one another. A flat text output can lose the relationship between those fields.

ADE Gen2 represents the document as a hierarchy. The document contains pages, pages contain blocks, and blocks may be text, tables, table cells, figures, signatures, stamps, logos or QR codes. This gives an automation system more context when deciding what each piece of information means.

The release includes two parsing models. DPT-3 Verity is designed for digitally created documents and returns a confidence score for every word. DPT-3 Pro is designed for more complicated files, including scanned pages, handwriting, non-Latin scripts, signatures and complex layouts. According to LandingAI, Verity uses roughly 40% of the credits charged by Pro, although you should test the result with your own documents before choosing a workflow. Source: MarkTechPost

The other important feature is atomic grounding. Each extracted item can be connected to a specific line or word on a page. For Verity, the system provides word-level grounding and a confidence value from 0 to 1. That means a reviewer can be sent directly to the uncertain word instead of checking the entire document.

The practical shift is simple: document automation should not only give you an answer; it should show you the evidence behind the answer.

How This Applies to Malaysian SMEs

1. Faster invoice and purchase order processing. If you operate a trading business, restaurant, workshop, distributor or construction company, your team may receive invoices in different formats. One supplier sends a clean PDF, another sends a photo taken with a phone, and a third uses a document containing tables with merged cells. A structured extraction system can identify supplier details, invoice numbers, dates, tax information, totals and line items before sending them into your accounting or approval workflow.

The grounding feature is especially useful when the extracted total does not match your purchase order. Instead of asking staff to search through a long PDF, your system can point them to the exact page and field used for the amount. This creates a clearer review process for finance staff and business owners who need to approve exceptions quickly.

2. Better handling of delivery orders and stock records. Malaysian SMEs commonly rely on delivery orders, goods received notes and handwritten signatures to confirm that stock arrived. A document system that recognises tables, signatures, stamps and item rows can help connect a delivery document to a purchase order and stock entry. If a quantity is unclear because of a blurred scan, the confidence score can route that item for manual checking instead of allowing uncertain data into your inventory record.

This is useful for businesses with multiple branches or warehouses. Your team can retain the original document while creating a searchable digital record of supplier, item, quantity, date and receiving staff. You can then investigate missing stock or delivery disputes with a specific visual reference rather than relying only on manually typed notes.

3. More reliable customer onboarding and compliance checks. Professional firms, recruitment agencies, education providers, property businesses and service companies often collect identity documents, application forms and signed declarations. These documents may contain personal information, so accuracy and traceability matter. Word-level grounding can support a review screen where staff see the extracted name, identification number or address alongside the original image.

Coordinate-based grounding may also support targeted redaction. For example, a workflow could identify and hide a specific identification number before a document is shared with another department. You still need proper access controls, retention rules and legal review, but precise document locations make those controls easier to implement than broad, approximate text masking.

4. Easier contract and service agreement review. If you manage maintenance contracts, rental agreements, supplier terms or customer service agreements, the important information may be spread across clauses, tables and signature sections. A structured document tree can help identify renewal dates, notice periods, service levels and signed pages. The system can then send reminders to a responsible person while linking each extracted value back to the relevant clause.

What the Numbers Tell You

The pricing model described for ADE Gen2 combines a page component with an output-character component. That matters because a short, simple document and a dense, information-heavy document may not consume the same amount of processing capacity.

Area Reported detail Why it matters to you
DPT-3 Pro 1 credit per page plus 0.5 credits per 1,000 output characters on priority processing Useful for complex scans, handwriting and difficult layouts
DPT-3 Verity 0.3 credits per page plus 0.2 credits per 1,000 output characters on priority processing Suitable for high-volume digital documents
Standard processing Reported at half the priority rates Suitable for background workflows that do not require an immediate result
Example workload A 12-page Pro parse returning 48,120 characters is reported at 36.1 priority credits Shows why document length and output volume both matter
Confidence scoring Verity returns a value from 0 to 1 for each word Supports automatic routing of uncertain fields to a human reviewer

All figures in the table come from the source article. Source: MarkTechPost Treat vendor-reported performance and workload estimates as starting points. Your results will depend on document quality, language, handwriting, layouts and the amount of information you ask the system to return.

Practical Takeaways

  • Start with one repetitive document type. Choose invoices, delivery orders or application forms rather than trying to automate every document at once.
  • Collect real samples. Include clear PDFs, phone photographs, scanned documents, handwritten pages and files from different suppliers.
  • Define the fields you actually need. Extracting every word may create unnecessary review work. Begin with supplier name, document number, date, totals and line items.
  • Set a human review threshold. Send low-confidence names, quantities, dates and totals to a staff member before updating your main system.
  • Keep the source reference. Store the page number, field location or evidence link with every important extracted value.
  • Separate urgent and background work. Use immediate processing when a customer or staff member is waiting. Run batch documents asynchronously when minutes or hours are acceptable.
  • Check privacy requirements. Review where documents are stored, who can access them and how long they are retained, especially for identity documents and payroll records.
  • Plan for integration. Confirm that the extraction output can connect with your accounting, inventory, CRM or approval tools.
  • Test migration carefully. Gen1 client code will not run against Gen2 endpoints, so technical changes are required when moving from the earlier version. Source: MarkTechPost

The Bigger Picture

For SMEs, the long-term importance of this trend is not that software can read a document. Optical character recognition has been available for years. The bigger development is that extraction systems are becoming more accountable and useful inside business processes.

A future workflow will not simply copy text from an invoice into a spreadsheet. It will identify the document type, extract selected fields, compare them with existing records, flag uncertain values, request approval and preserve the evidence used for each decision. That is much closer to how a careful administrator works.

This also changes how you should evaluate automation tools. Ask whether the system can handle your actual documents, not just a clean demonstration file. Ask whether it can identify uncertain results, preserve the original evidence and connect with the software your team already uses. Ask what happens when the document contains a signature, a stamp, a table with merged cells or a mixture of English and other scripts.

For a Malaysian SME, the best first step is a controlled pilot. Choose a document process that consumes staff time every week, measure how often fields are extracted correctly, record the exceptions and calculate how much manual checking remains. If the system provides reliable evidence and fits your approval process, you can gradually extend it to other departments.

Document extraction is becoming less about reading files and more about building dependable business records. When you can see both the extracted answer and the exact source behind it, your team can move faster without giving up review and accountability.

Ready to Streamline Your Operations?

Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →