Skip to main content

40 docs tagged with "OCR"

View All Tags

Activate and renew OCR license

The Work.AI Developer installation requires a separate OCR license that is different from the server installation of Work.AI. Developer comes with two basic licenses for default ABBYY FRE 11 and FRE 12 limited by time and the number of pages used. The out-of-the-box OCR license expires three months after the Work.AI Developer release or reaching the default 1,000 pages.

Calculate capacity

To plan capacity means to ensure the availability of a cost-effective capacity that meets or exceeds the business needs as established in Service Level Agreements (SLAs) at any given time. Within the context, capacity is defined as the maximum load the Work.AI platform can handle.

Configure advanced OCR settings

The guide describes the advanced configuration of the Optical Character Recognition (OCR) feature in Work.AI. In most cases, the default configuration is enough, but if you need fine-tuning for your specific case, see the instructions below.

Create and manage templates

Templates are intended to simplify document ingestion by enabling automated information extraction from structured forms (for instance, ACORD). When a template is created, blank forms are annotated with labels. Based on these labels, data is then extracted from input documents.

Data Analyst workflow

A Data Analyst (DA) acts based on several factors originating from the customer's data analysis. The article shares a typical workflow for new DAs.

Disable OCR workers

In the WorkFusion platform, the OCR component is installed by default. If you do not plan to use OCR, you can disable it. This will increase the throughput for Control Tower, RPA, and AutoML workers.

Extract machine-readable zone

The machine-readable zone (MRZ) is typically found on official travel or identity documents of many countries. It can have two or three lines of machine-readable data. A new profile allows processing MRZ written in accordance with ICAO Document 9303.

Manage datasets

A dataset is a container for documents and all related meta information, such as:

OCR

High-level description

OCR

The article dwells on the obsolete approach of OCR usage in the scope of the ODF 2 framework.

OCR analysis and tuning

The goal of this stage is to achieve acceptable OCR configuration and get understanding of final OCR quality.

OCR plugins

The ocr plugin works in the synchronous mode, and therefore, becomes a performance bottleneck if used within a BP where the OCR page volume is high, or documents contain more than 20 pages. The recommendation is to use the prepared BP artifacts where the OCR logic is asynchronous.

Optimize OCR process

A standard WorkFusion OCR use case provides a set of Bot Configurations that can be re-used or extended to address custom OCR challenges.

Recognize Windows checkmarks

The checkmark recognition feature is available for the WorkFusion platform up to v10.1.6 and IA Cloud Developer/Work.AI Developer (all versions).

Split into batches

Goal of the step: Split documents into batches by layout types before starting any manual handling—labeling, reviewing, and so on.

Study documents

Before labeling, the most essential information the subject-matter experts (SME) should share with the data analyst (DA) includes the following:

Use OCR Bridge step

The article is an example of the OCR Bridge step usage in the scope of the ODF 2 framework. The Bot Tasks used in the example are a part of Processing Business Process from the example project.