Skip to main content
Version: 10.2.9

Document labeling and validation

Goal: get labeled dataset of high quality.

Input: Manual Task, batches of documents.

Output: reviewed data categorized by groups based on the type of inconsistency found.

Labeling

Once qualification is finished, it’s time to go ahead with document tagging.

To start labeling, each SME should be provided with:

  • A link to a task in Workspace
  • A file with general information about labeling and Workspace tools
  • Specific features and labeling logic of this particular task.

Here is an example of an email that a Data Analyst (DA) can send to the team of SMEs and Delivery Manager.

Hi (SME's_name),

The Manual Task is ready and available for data labeling.

The name of the task is “Tagging_Georgia_batch_1” with 64 documents in it. Here is the link to the task: https://wf-app.workfusion.com/workfusion/login

Attached you’ll find a file with instructions. Please, carefully review the sheet “Specific instructions” with the labeling logic description and, if necessary, look through “General instructions” to stay aligned with basic labeling rules.

In case of any questions, feel free to ask.

Best regards,

When a Subject Matter Expert (SME) finished labeling, the DA should provide him or her with another batch. If a new batch includes documents of another layout, a new file with instructions should also be sent.

Labeling report

By the end of the day, the DA should provide a report about labeling progress to all the interested parties, but mainly to the Delivery Manager, the SMEs’ manager and SMEs.

The following template can be used as a report:

You may create your own labeling tracker, but it should include the following information:

  • Name of SME
  • Number of tasks performed for each day
  • Average speed per task
  • Comments on most common mistakes
  • Overall documents labeled
  • Number of documents left

Dataset validation

When at least one batch is finished, the DA can start validation. For demo purposes, we used validation via browser:

  1. Download a snapshot from the completed Manual Task.
  2. Upload it into the Save labeled content to S3 Business Process. The output will be the same CSV file but with links to the labeled data (HTML) instead of XML content.
  3. Insert an additional column to enable categorization. For more information about category drop-downs, see Labeling and validation.
  4. Create an Excel copy of the CSV file to keep all formatting.
  5. Open each document in a browser and make sure:
    • All the required fields are labeled. Make sure the n/a option is soundly chosen, and no value is missed due to inattentiveness, except for the cases where it's totally corrupted by OCR.
    • If a value is present in the document, it is labeled. Otherwise, missed fields will be badly represented, and such documents should be re-labeled or the dataset should be increased.
    • Fields are labeled from the same location across the entire dataset. If more than one labeling option for the same field is found in one template, see how often this happens. If, in two out of 50 documents, registered_agent_name is labeled in some other place, it’s better to exclude such documents.
    • Pay attention to document layouts as well. Note that due to OCR, the layouts of one vendor can be multiplied, for example, because of a broken table structure, and so on. Such layouts and badly represented layouts (less than 10 for one vendor) should be increased, if possible.
    • The value is labelled entirely.
    • Data value wasn't changed manually. For example, the labeled string for entity_type is Limited Liability Company, the data value should also be Limited Liability Company, not LLC.
    • There is no labeling of corrupted values that cannot be restored.
    • The Append selected to option was used correctly. If you are not sure, just try to select a desired value. If no extra information (except for the required text chunk) gets into your selection, there are no appended parts.
  6. Depending on the type of the mistake found (if any), assign one of the categories to each record: Good, Bad OCR or Re-label.

In the end, your report should look like this.

To get more details, download the Excel file.

If we take a "Georgia" layout as an example, we'll see that out of 68 documents of the first batch, 23 have different types of inconsistencies and can be included neither in training nor in a test set without correction, and four records within category "Bad OCR."

The same validation logic should be applied to the rest of the batches for all templates. So, the number of revealed inconsistencies for the rest of the states is the following: New Hampshire, 36 documents (out of which five had bad OCR); Massachusetts, 29; California, 35 (8 with bad OCR); and the other two batches of Georgia, 11. Total number of records to be corrected: 121 (23+31+29+27+11).

All the incorrect documents in the "Re-label" category, collected from all the batches will be an input to a separate Manual Task for re-labeling. As a result, we end up with 593 absolutely correct documents.