Skip to main content

Issues related to input data

The article addresses OCR issues associated with quantity or quality of input documents.

Increased file quantity

To investigate whether the cause of your issue is increased file quantity, perform the following steps:

  1. Compare the number of processed files at good performance indicators and at poor ones.

  2. Based on the comparison, if the number of files has increased, launch the task with the previous file quantity.

  3. Check whether the same file quantity was covered by load testing.

Large input file

To investigate whether your issue is due to the input file having excessive size, perform the following steps:

  1. Locate the input file you suspect might be causing trouble, compare it with a document without the same problem. Pay attention to such characteristics as the number of pages or tables and so on.

  2. Try reproducing the issue with the document on a different environment (for instance, dev or test one) with the same configuration.

  3. Find out whether the document was covered by testing.

Document quality

To investigate whether your issue is due to the quality of an input document, perform the following steps:

  1. Check the OCR step execution timestamp in the framework Business Process (BP). If you can see that the document causes the process to start slower, locate the document using its UUID. Then, download it to look through it manually. If it is in the cloud, open it in a browser.

  2. If no timestamps exist, add them to each document manually and rerun the BP in the current environment or locally. Find the documents processed longer than the average processing time.

info

Do not forget to delete the timestamps after the check if you added them manually.

Output

If you confirm any of the issues, escalate them to the Support team. Otherwise, continue the investigation.

View also: