Skip to main content
Version: 10.3

Documents labeling and validation

Goal: Get labeled dataset of high quality.

Input: Manual Task, batches of documents.

Output: Reviewed data categorized by groups based on the type of inconsistency found.

Labeling

Once qualification is finished, it’s time to go ahead with document labeling.

To start labeling, each SME should be provided with:

  • A link to a task in Workspace
  • A file with general information about labeling and Workspace tools
  • Specific features and labeling logic of this particular task.

Here is an example of an email a DA can send to the team of SMEs and Delivery Manager.

Hi (SME's_name),

Manual Task is ready and available for data labeling.

The name of the task is “Tagging_Georgia_batch_1” with 64 documents in it. Here is the link to the task: https://wf-app.workfusion.com/workfusion/login

Attached you’ll find a file with instructions. Please carefully review the sheet “Specific instructions” with labeling logic description and, if necessary, look through “General instructions” to stay aligned with basic labeling rules.

In case of any questions, feel free to ask.

Best regards,

When an SME has finished labeling, the DA should provide him or her with another batch. If a new batch includes documents of another layout, a new file with instructions should also be sent.

Labeling report

By the end of the day, a DA should provide a report about labeling progress to all the interested parties, but mainly to the Delivery Manager, the SMEs’ manager and to SMEs.

The following template can be used as a report:

You may create your own labeling tracker, but it should include the following information:

  • Name of SME
  • Number of tasks performed for each day
  • Average speed per task
  • Comments on most common mistakes
  • Overall documents labeled
  • Number of documents left

Dataset validation

When at least one batch is finished, DA can start validation. For demo purposes, we used a method of validation via browser:

  1. Download snapshot from the completed Manual Task
  2. Upload it into the Business Process "Save labeled content to S3." Output will be the same CSV file but with links to the labeled data (HTML) instead of XML content. 
  3. Insert an additional column to enable categorization. For more information about how to enable category drop-downs, see labeling and validation.
  4. Create an Excel copy of CSV file to keep all formatting.
  5. Open each document in browser and make sure:
    • All the required fields are labeled. Here you need to make sure the "n/a" option is soundly chosen and the value is not missed due to inattentiveness, except for the cases where it's totally corrupted by OCR.
    • If the value is presented in the document, it was labeled. Otherwise missed fields will be badly represented and such documents should be re-labeled or the dataset should be increased.
    • Fields are labeled from the same location across the whole dataset. If more than one variant of labeling of the same field is found in one template, see how often this happens. If two out of 50 documents in registered_agent_name was labeled in some other place, it’s better to exclude such documents.
    • Pay attention to documents’ layouts as well. Note that due to OCR, the layouts of one vendor can be multiplied — for example, because of a broken table structure, etc. Such layouts as well as badly represented layouts (less than 10 for one vendor) should be increased, if possible.
    • Complete value was labeled.
    • Data value wasn't changed manually — for example, the labeled string for entity_type is "Limited Liability Company," the data value should also be "Limited Liability Company, " not "LLC."
    • There is no labeling of corrupted values which cannot be restored.
    • Option “Append selected to” was used correctly. If you are not sure, just try to select a desired value. If no extra information (except for the required text chunk) gets into your selection, there are no appended parts.
  6. Depending on the type of mistake found (if any), assign one of the categories to each record: Good, Bad OCR or Re-label (for more information about categories, click  here ).

In the end, your report should look like this.

If we take a "Georgia" layout as an example, we'll see that out of 68 documents of the first batch, 23 have different types of inconsistencies and can be included neither in training nor in test set without correction, and four records within category "Bad OCR."

The same validation logic should be applied to the rest of the batches for all templates. So the number of revealed inconsistencies for the rest of the states is the following: New Hampshire, 36 documents (out of which five had bad OCR); Massachusetts, 29; California, 35 (8 with bad OCR); and the other two batches of Georgia, 11. Total number of records to be corrected: 121 (23+31+29+27+11).

All the incorrect documents — in category "Re-label" — collected from all the batches will be an input for separate Manual Task for re-labeling. As the result, we end up with 593 absolutely correct documents.