Documents labeling and validation
Goal: Get tagged data set of high quality.
Input: Manual Task, batches of documents.
Output: Reviewed data categorized by groups based on the type of inconsistency found.
Labeling
Once qualification is finished, it’s time to go ahead with document tagging.
To start tagging, each SME should be provided with:
- a link to the task on WorkSpace
- a file with general information about tagging and WorkSpace tools
- specific features/tagging logic of this particular task.
Here is an example of an email a DA can send to the team of SMEs and Delivery Manager.
Hi (SME's_name),
Manual task is ready and available for data tagging.
The name of the task is “Tagging_Georgia_batch_1” with 64 documents in it. Here is the link to the task: https://wf-app.workfusion.com/workfusion/login
Attached you’ll find a file with instructions. Please carefully review the sheet “Specific instructions” with tagging logic description and, if necessary, look through “General instructions” to stay aligned with basic tagging rules.
In case of any questions, feel free to ask.
Best regards,
When an SME has finished tagging, the DA should provide him or her with another batch. If a new batch includes documents of another layout, a new file with instructions should also be sent.
Labeling Report
By the end of the day, a DA should provide a report about tagging progress to all the interested parties, but mainly to the Delivery Manager, the SMEs’ manager and to SMEs.
The following template can be used as a report:
You may create your own tagging tracker, but it should include the following information:
- Name of SME
- Number of tasks performed for each day
- Average speed per task
- Comments on most common mistakes
- Overall documents tagged
- Number of documents left
Data Set Validation
When at least one batch is finished, DA can start validation. For demo purposes, we used a method of validation via browser:
- Download snapshot from the completed Manual Task
- Upload it into the business process " Save tagged content to S3." Output of this process will be the same CSV file but with links to the tagged data (HTML) instead of XML content.
- Insert an additional column to enable categorization. For more information about how to enable category drop-downs, see labeling and validation.
- Create an Excel copy of CSV file to keep all formatting.
- Open each document in browser and make sure:
- All the required fields are tagged. Here you need to make sure the "n/a" option is soundly chosen and the value is not missed due to inattentiveness, except for the cases where it's totally corrupted by OCR.
- If the value is presented in the document, it was tagged. Otherwise missed fields will be badly represented and such documents should be re-tagged or the data set should be increased.
- Fields are tagged from the same location across the whole data set. If more than one variant of tagging of the same field is found in one template, see how often this happens. If two out of 50 documents in registered_agent_name was tagged in some other place, it’s better to exclude such documents.
- Pay attention to documents’ layouts as well. Note that due to OCR, the layouts of one vendor can be multiplied — for example, because of a broken table structure, etc. Such layouts as well as badly represented layouts (less than 10 for one vendor) should be increased, if possible.
- Complete value was tagged.
- Data value wasn't changed manually — for example, the tagged string for entity_type is "Limited Liability Company," the data value should also be "Limited Liability Company, " not "LLC."
- There is no tagging of corrupted values which cannot be restored.
- Option “Append selected to” was used correctly. If you are not sure, just try to select a desired value. If no extra information (except for the required text chunk) gets into your selection, there are no appended parts.
- Depending on the type of mistake found (if any), assign one of the categories to each record: Good, Bad OCR or Re-Tag (for more information about categories, click here ).
In the end, your report should look like this.

If we take a "Georgia" layout as an example, we'll see that out of 68 documents of the first batch, 23 have different types of inconsistencies and can be included neither in training nor in test set without correction, and four records within category "Bad OCR."
The same validation logic should be applied to the rest of the batches for all templates. So the number of revealed inconsistencies for the rest of the states is the following: New Hampshire, 36 documents (out of which five had bad OCR); Massachusetts, 29; California, 35 (8 with bad OCR); and the other two batches of Georgia, 11. Total number of records to be corrected: 121 (23+31+29+27+11).
All the incorrect documents — in category "Re-Tag" — collected from all the batches will be an input for separate Manual Task for re-tagging. As the result, we end up with 593 absolutely correct documents.
Also, see Final dataset review.