Document Extraction

Automatically extract data from physical and digital documents using OCR and artificial intelligence.

How does it work?

The platform combines OCR technology (optical character recognition) with artificial intelligence models to extract structured data from documents such as invoices, receipts, contracts, identity documents, and printed forms. You upload an image or PDF and the platform automatically identifies and extracts the fields you need.

Upload a document

  1. Go to the Documents section and click "Upload document".
  2. Drag or select the file. Supported formats: PDF, PNG, JPEG, WEBP.
  3. Select the extraction template that defines which fields to extract (or leave automatic extraction).
  4. The platform will process the document and display the extracted data for your verification.

Extraction process

The document goes through several stages:

  • Initial OCR — All text, tables, and fields in the document are detected.
  • Intelligent extraction — AI maps the detected text to the fields defined in the template.
  • Confidence score — Each extracted field receives a confidence score (0-100%).
  • AI fallback — If confidence is low or required fields are missing, a more advanced AI model analyzes the document.

Supported documents

  • Invoices and payment receipts
  • Identity documents (IDs, passports)
  • Contracts and legal documents
  • Printed forms
  • Purchase orders
  • Word documents (.docx) — AI analysis

Siguiente paso

Learn how to create extraction templates to define which data to extract from each document type.

Documentation | Gloo Platform by SYNOV