Skip to main content
Version: 10.3.1

Labeling requirements

In order to ensure a dataset is of high quality, result of labeling must be:

  • Consistent
  • Complete
  • Normalized
  • Diverse
  • Have Correct Data Values
  • Comprehensive

Consistency

tip

Each field in the dataset should be labeled consistently for ML training, i.e. fields must be labeled in the same place across layout. The occurrence of the field in the document has to be chosen and adhered to across all documents in the batch.

Expand to see the example

There are three possible ways to extract the field business_id, its value is C 133395, and there are 3 occurrences of this field. Let’s try to describe the value itself and its context for every occurrence:

Case 1

Letter and digits after “Annual report for” in the first string before the table

Case 2

Letter and digits after “No.” in the left upper corner cell of the table

Case 3

Letter and digits in a cell of the table after: “5. Organized under the Laws of: ID”.

Case 1. Correct

If we label this field in the same place all the time (let’s choose the second case, for example) we will get the same description for every document in our training set. So, after the training, model will be 100% confident that for the field business_id needs to be extracted. That is a letter and digits after “No.” in the left upper corner cell of the table.

Case 2. Incorrect

If we label this field in different places all the time (let’s suppose we label in the same proportion), after the training, model will be:

  • 33% confident that for the field business_id needs to be extracted. Letter and digits after “Annual report for” in the first string before the table.
  • 33% confident that for the field business_id needs to be extracted. Letter and digits after “No.” in the left upper corner cell of the table.
  • 33% confident that for the field business_id needs to be extracted. Letter and digits after “5. Organized under the Laws of: ID” in a cell of the table.
note

The model does not treat text same way people do. While people can perceive text entirety, the model perceives text as strings of symbols organized in an HTML or XML tree. The goal is to identify the value and context surrounding the value.

Expand to see read more

Context is basically everything that surrounds the labeled piece of text. It is the most important source of features and is used to determine where value should be searched for.

The set of features is defined on the whole dataset (all the documents). In each document, for each gold value feature, values will be calculated and weighed during the training process and the model (the function) will be built. When the model is trained and you run it on some data, it will extract the strings that can be classified as "the value should be extracted" by their features. So Feature represents a simple question such as "Is the current cell located in the first column?" with a clear non-ambiguous answer "Yes/No" or "1/0". Learn more about features here.

Values themselves are the source of features as well. For example, CUSIP number that is used for classifying financial instruments worldwide has a definite shape and check sum, so any extracted value for CUSIP can be validated by these criteria.

Wrong values or wrong context are the source of incorrect features and weights, which affects the model's quality. If the value is not extracted in the context the model expects it to be, it may learn that the values in this context may be incorrect and should not be extracted at all. So It will generate additional FNs in extraction results. Another bad effect may take place if the model defined the features for context in not an optimal way. For example, in one type, invoice_date is given in the document after "Invoice date:" and after "Date:" and is labeled in both contexts. Documents also contain "Payment date: {date in the same format}". In this case, the model may find "the context like required" and extract payment date instead of invoice date and has incorrect feature "date: {value that should be extracted}". So it will generate additional FPs in extraction results.

Completion

Completeness has two different aspects:

  • If a correct value in the correct form and with the correct context is given in the document, it must be labeled.
  • If a field value contains multiple words, we should always label the whole value with all the words belonging to the field, not only a part of it.
note

In case, it's necessary to exclude some part from the values (for example, there is a requirement that supplier_name shouldn't contain legal endings), it's recommended to label the full value, correct the data value and apply post-processing to the model results, so that you have full control of what is changed.

Normalization

Data value can be given in many different formats across documents. It should be normalized to the same one format in all cases for several reasons: For some fields (for example dates or amounts), the values are normalized in a manual task. They should be normalized accordingly after ML extraction. Otherwise, it will be difficult to count the statistics correctly. Another reason for normalization is that the customer may need to upload this data in SAP or some data base and all the values should be of some particular format.

FieldIn documentsNormalized
Date01/18/2017, January 18, 2017, Jan-18-201701/18/17
Amount10,000,000.00 or 10,000,000 or 10 000 000.00 or 10 000 00010000000.00
Company NameWorkFusion or WorkFusion, LTD. or WFWorkFusion

Diversity

A dataset should have good representation of all layouts and fields so as to train ML well. Objects from bad-represented or not-represented layouts may have feature values which make the model treat them like outliers.

The required number of documents to train a layout/field is not fewer than 20–50 examples. Also note that to launch the training you need to use not fewer than 20 documents, even if there is only one template in the dataset.

Correct data values

Correctness of data values means that in each document, for every configured field, should present correct data value. In other words, the values should be exactly the same as we want the ML model to extract from the original document. Usually, mistakes are caused by:

  • OCR, like iOO8.4O instead of 1008.40 (data value should be corrected manually in the task),
  • By user (if the value is wrongly corrected manually or the wrong value is selected from the list).

Data value is the information which is given in the answer box. It can differ from the labeled piece of information when formatting mask is applied.

Comprehensiveness

The value and especially context should be comprehensive and accurate so as to provide correct features and weights for extraction by model. So in this regard, attention should be paid to OCR quality of each field and its context.

The following cases should not be labeled and provided for model training.

OCR issues

  1. OCR has not recognized the field at all.

    Change OCR parameters. If this doesn't help, exclude document from training set.

  2. OCR has corrupted the value.

    If the value itself and its context are totally corrupted and without checking the original document you cannot understand what should be in this part of a document, it should not be labeled. The document should be excluded from the training set, because it won't let the model learn anything useful and may be the source of wrong features and weights. If such documents are regular cases in the common documents flow and are not caused by the original documents quality, it's recommended to improve OCR quality.

  3. OCR has not recognized the context at all.

    If the context of the field is missing or corrupted completely, it will be difficult for the model to recognize a required value and extract it with high confidence, so such values should not be labeled, "N/A" option should be chosen. Write a comment about bad OCR to be able to find such records later.

  4. OCR has corrupted the context.

    Solution is the same as above.

  5. OCR has corrupted the structure, only the part of string is field's value.

    Original document

    After OCR

    Text may be converted into a tables with random structure that won't be reproduced in other documents. If the value to be extracted is in the table, it's recommended to exclude such documents from the training set as they may be a source of wrong features and weights. If such documents are rather regular cases in the common documents flow and are not caused by the original documents quality, it's recommended to improve OCR quality.

    If doesn't help, exclude document from training set.

The reason for not labeling values with broken structure with the help of "Append selected" is that model reads entire string: "FULL SPECTRUM GARDEN 35 ROSEMARY LANE, 001182906".

From a human perspective, we can define the borders and ending of the value, but we can't teach a model well where to stop extracting if we have the value in different strings surrounded by extra data. The model might learn well if the value is constant (for example, there are 30 more documents of this supplier and in all of them, the value is split in that way). But the most likely case is we'll receive a FN or FP/FN — the model won't extract anything or will extract only part of the value. But we can also receive the following cases in other batches: