Skip to main content
Version: 10.2.9

Get qualitative OCR result

Goal:

  • Get a high-quality dataset ready for labeling.
  • Understand the layout representation (for instance, the number of documents for each layout) to collect a good training set and a representative test set.

Input: CSV file with links to original documents. All documents are combined with no sorting by layouts.

Output: OCRed dataset in which at least 90% of documents have acceptable quality.

To generate links to original documents, follow the steps below:

  1. Open the S3 browser.
  2. Choose a bucket (or create one) and create a separate folder to store original documents.
  3. Click the Upload button and choose Upload Folder(-s).
  4. Select a folder on your computer with original documents and wait for all records to be uploaded.
  5. Set the permission to let all users read the files:
    1. Select all files.
    2. Click the Permissions tab.
    3. Select Read next to All Users.
    4. Click Apply changes.
  6. Select all records in the folder (Ctrl+A), right-click the selected files, and then click Generate Web URL(s).
  7. In the pop-up window, click Copy to clipboard and verify that all links were copied.
  8. Open a new Excel file and paste the links.

Validate initial OCR quality

Before sending all documents to OCR, make sure that current OCR settings are sufficient to ensure 90%+ quality of the dataset after OCR. To assess the OCR quality, prepare a sample set: randomly choose 20% of the dataset documents and OCR them.

Analyze results

When analysis of each record is completed, calculate the percentage of documents with acceptable quality. For more information, see OCR analysis and tuning.

When removing documents with unacceptable quality, pay attention to what template type they belong to. If unacceptable quality items are mostly documents of the same template, you might need to request additional documents. In this situation, notify the Delivery Manager.