Paper Type

ERF

Abstract

Modern enterprises maintain extensive repositories of image-based business documents within intranet systems, creating a critical need for automated processing to enhance efficiency. Unlike standardized forms, most business documents are semi-structured, with layouts and field positions varying across organizations and document types. This variability constrains conventional Optical Character Recognition (OCR), which relies on location-based extraction, and Key Information Extraction (KIE) techniques that require costly domain-specific pre-training for new formats. We propose a Vision Language Model (VLM)-based process that treats documents as image inputs and jointly models visual layout and textual semantics to extract and organize key elements. By leveraging semantic content and spatial cues, the framework improves extraction accuracy and economic efficiency. Experimental results show that the VLM-based approach significantly outperforms existing OCR and KIE solutions in both precision and operational cost. A human-in-the-loop mechanism enhances reliability, providing a practical framework for semi-structured document automation in dynamic enterprise environments.

Paper Number

1824

Comments

AI SYSTEM

Share

COinS
 
Aug 15th, 12:00 AM

Beyond OCR: Leveraging Vision-Language Models for Semi-Structured Business Documents

Modern enterprises maintain extensive repositories of image-based business documents within intranet systems, creating a critical need for automated processing to enhance efficiency. Unlike standardized forms, most business documents are semi-structured, with layouts and field positions varying across organizations and document types. This variability constrains conventional Optical Character Recognition (OCR), which relies on location-based extraction, and Key Information Extraction (KIE) techniques that require costly domain-specific pre-training for new formats. We propose a Vision Language Model (VLM)-based process that treats documents as image inputs and jointly models visual layout and textual semantics to extract and organize key elements. By leveraging semantic content and spatial cues, the framework improves extraction accuracy and economic efficiency. Experimental results show that the VLM-based approach significantly outperforms existing OCR and KIE solutions in both precision and operational cost. A human-in-the-loop mechanism enhances reliability, providing a practical framework for semi-structured document automation in dynamic enterprise environments.

When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.