Skip to main content
Version: 10.3.1

Prepare for labeling

Goal: Train SMEs to label the documents properly to ensure high-quality dataset and to prepare documents for labeling

Input: Original documents in XML/HTML format, SMEs team

Output:

  • Documents prepared for labeling: split into batches, with labeling logic predefined.
  • SMEs trained to label documents qualitatively.

This stage is critically important, because here the Data Analyst lays the foundation for the whole dataset collection process. If there are any issues at this stage not revealed and addressed properly, they can result in major problems during later stages.

The main aim of this stage is to prepare everything for labeling to start: to train and qualify SMEs, and to prepare documents for labeling. SMEs' training and qualifying means that they are able to label documents and meet the requirements of a high-quality dataset and documents preparation, including splitting into batches, establishing labeling logic and describing it in labeling instructions.

DA needs to perform the following steps in this stage:

  1. Manual task design
  2. Train SMEs
  3. Split into batches 
  4. Labeling logic definition
  5. Labeling instructions provision

In the end of this stage, the Data Analyst assigns a specific manual task for each SME, provides a batch and instructions for it, and creates a data labeling tracker to track SMEs' speed of labeling.

In the result of this stage, the batches and manual tasks are prepared and assigned to SMEs, a report on the whole stage is provided, and all the necessary preparation for the first labeling iteration is conducted.

The following template can be downloaded and used as a report for this stage:

Risks at the stage:

  • Wrongly configured Manual Task

    The Manual Task should be configured according to business logic of all the documents. If a mistake is revealed in the process of labeling, it will take extra time and effort to reconfigure the manual task, re-label documents, and even retrain the model (if training iterations start in parallel with labeling).

  • Low labeling quality

    There is a risk to involve under-skilled SMEs for labeling, which could result in inconsistent and incorrect labeling, hence, bad quality for the dataset or extra time to re-label and verify documents. SMEs must be trained well or at least pass through a qualification task after additional training.