Introduction
How do I set up my task to leverage Virtual Data Scientist?
The Virtual Data Scientist applies numerous machine learning algorithms to each task to determine the best automation method.
Some tasks naturally fit certain categories of automation approaches, and VDS is optimized for this. Generally speaking, one category of automation matches each human interface template.
For the processing of text (web data, news articles, short samples of text, rendered OCR, and so on), the following task types are applicable for the Virtual Data Scientist:
Information Extraction extracts key phrases from natural language documents by identifying patterns.
Categorization assigns the natural language document to one or more categories within an ontology (fixed, hierarchical set of categories).
Topicality generates a list of prominent subject matters from the natural language document.
Augmentation interprets and annotates natural language documents with the knowledge from external sources.
For the processing of images (photographs, smart-phone images, scanned documents, and so on), applicable task types are as follows:
Object Detection identifies, counts, and analyzes entities within the image by creating a neural representation of the various types of objects that are known in our database.
Object Classification assigns a categorization and attributes to entities within the image, such as product identification or people recognition.
For the processing of non-digital or scanned documents, the applicable task type is Optical Character Recognition (OCR). The task automatically converts any document, PDF, image, or handwritten content into a computer-searchable and readable text, preserving the original layout. WorkFusion learns spatial clues and styles to correct potential spelling errors, recognize uncommon characters, and improve low-quality reproductions or complex content.
What types of tasks are applicable for Virtual Data Scientist?
That’s the biggest advantage of the Virtual Data Scientist. The VDS will automatically run experiments and build machine solutions for any human task that fits into one of the VDS-configured task templates.
How is the automation rate estimated?
The Virtual Data Scientist performance is measured directly by scoring (testing) the performance of the algorithms against a small but statistically representative set of human completed work. This procedure is a part of Statistical Quality Control to evaluate and tune algorithms and ensure that all reported statistics are calibrated and precise.
Typically, for tasks that are VDS-configured and well suited for machine learning, the Virtual Data Scientist begins reporting some automation after several hundred training examples. See How long will it take to see results? This number varies significantly depending on the complexity of the source and the desired products.
Very roughly, doubling the number of training examples will drop the error rate by about 10%, plateauing at some point. After approximately a thousand training examples, the Virtual Data Scientist determines whether the task can be automated and what performance can be projected.
How long will it take to see results?
Typically, the Virtual Data Scientist can begin reporting early results after several hundreds of training examples. This number varies significantly. What is important is that, after approximately a thousand of training examples, the Virtual Data Scientist will determine whether the task can be automated and what performance can be projected.
For a text, this number also varies significantly, for instance, depending on the number of categories to be categorized, the length of a document, the number of fields to be extracted, whether the document is a web text or standardized text. For images, an order of magnitude (10x) more training examples are needed, and the clarity and conditions of the image matter greatly.
How do I see my automation results?
Automation results can be viewed by going to the View Results section of a Business Process. This has been described in the View Automation Results section.
Why automation rate was different from actual one (higher or lower)?
The automation rate is a direct measurement of the algorithm performance on unseen and newly completed examples. The automation rate presented by the Virtual Data Scientist represents the performance up to a date. Projecting that performance forward assumes that future unseen examples will be of roughly the same style, structure, and content.
Occasionally, the sources of a document change over time, or styles fluctuate, which can cause a small (few to ten percent) drop in the automation rate as compared to the performance up to a date.
The Virtual Data Scientist automatically interprets signals from Statistical Quality Control (SQC) to retrain and tune automation algorithms as the performance fluctuates. Mind that SQC always delivers the assigned quality level even if the algorithm performance fluctuates.
How does Virtual Data Scientist work?
The process is analogous to teaching an apprentice by an example and testing periodically with a pop quiz. After each quiz, the lesson is revised to create new ways to explain the most important and most confused points to the apprentice. Once the apprentice demonstrates proficiency, he or she begins repeating the task on their own, periodically taught on newer and more challenging examples.
The Virtual Data Scientist sources from countless algorithms to transform each human response and document (or image) into machine-interpretable representations. The VDS trains (teaches) machine learning algorithms using these representations and tests. It can intelligently create and select new machine representations and configurations of algorithms to improve until producing a few comprehensive highest-performing automation algorithms.
The Virtual Data Scientist recommends these algorithms to the user. Once the process is automated, the Virtual Data Scientist receives signals from Statistical Quality Control to calibrate and optimize the algorithm and repeats the training process when Statistical Quality Control indicates that a higher performing model can be found.
What is SQC and how does it work?
Statistical Quality Control is an application of industrial process management and control methods for ensuring highly trustworthy production from an assembly line process. WorkFusion SQC methods choose an optimally cost-effective number of units to test for accuracy in a way that is statistically guaranteed to never allow a process to produce below its acceptable quality level.
SQC methods are designed to ensure the continuity of accurate processes knowing that the quality of its components will vary. Within WorkFusion, for example, SQC moderates for variations in worker quality and temporary fluctuations of machine performance as source material changes. These methods provide another powerful advantage of WorkFusion human-in-the-loop solutions over machine-only and human-only solutions.
SQC also provides the Virtual Data Scientist with signals to tune algorithms and calibrate all machine and worker analytics for accurate real-time monitoring of the process. This allows the Virtual Data Scientist to test and train machine learning algorithms without needing a separate test environment so that every human labor can be delivered to your customers.
What happens if Crowd answers questions wrong?
The Virtual Data Scientist trains algorithms to replicate the work of cloud workers. If workers consistently and overwhelmingly repeat the same wrong answers, the Virtual Data Scientist will learn these wrong answers.
Thankfully, Virtual Data Scientist algorithms are very robust against occasional worker mistakes and temporarily underperforming workers. The Virtual Data Scientist is the wisdom of a Crowd.
Occasional mistakes by high-accuracy workers are caught and corrected by the adjudication process. The Virtual Data Scientist identifies low-accuracy and underperforming workers and provides additional adjudication to their answers to ensure high overall accuracy before delivering to the customer or training an algorithm.
The Virtual Data Scientist understands the underlying structure of a text, images, and data and trains algorithms that downweigh the answers that do not match. For this reason, even many wrong answers from several workers can still result in a correctly trained model.
How does SQC work once I apply my recommendations?
SQC sends assignments to cloud workers to ensure an acceptable quality level is maintained and the measured machine and worker metrics are accurate. SQC continues this process after machine automation begins.
SQC chooses the optimally cost-effective combination of automated machines and cloud work that always delivers at or above an acceptable quality level.
Cloud workers statistically verify that the automated and delivered results match the predicted accuracy based on the historical performance. SQC will automatically increase and decrease the amount of work sent to cloud workers to maintain the process performance at the requested acceptable level. SQC adapts during periods of repeated overperformance of the algorithms by decreasing human cloud work. Repeated periods of algorithm underperformance signal the Virtual Data Scientist to retrain the algorithm.
Why does it show ML models are being trained but does not show them connected to all workflow aspects?
The Automation Business Process contains three sub-processes:
- Training
- Machine model and Crowd task for failed records
- Statistical Quality Control
When Crowd answers a preset number of records, these records are written to a Training Set. The Training sub-process uses the Training Set for model training. When the user applies the recommendation and runs the process, records first go through the machine-learned model and then to a Human Task (if the machine one failed) and SQC (if enabled).
The results of these sub-processes are saved in the same Training Set. The Training Business Processes will again use this Training Set. Hence, not all aspects of the workflow need to be directly connected to the Training sub-process. The training sub-process runs in the background and retrains the model as the Training Set gets more records.
How will I know which records have been SQCed and if SQC confirmed or rejected automation?
Records processed by cloud workers as part of Statistical Quality Control will be marked as indicated, including their target quality, inspection result, and batch ID. These records will also be marked with the current status of the SQC process and information on whether SQC has confirmed or rejected the automation.
If accuracy setting is 95%, how do I make sure the 5 incorrect percent of the records do not go to client?
Accuracy is the measurement of correctness across a collection of documents. Some documents will have been processed correctly, and a few will be incorrect. When the accuracy is set to 95%, approximately 5 out of 100 records will be incorrect. This can be remedied by increasing the accuracy setting to reduce the number of incorrect ones or to provide an additional layer of verification before sending it to your client. From within a single task, there is no way to know which examples are correct and which are not.
When Virtual Data Scientist reports an accuracy setting of 95%, does that reflect accuracy for fields or for records?
The Virtual Data Scientist can be configured to report and train based on several definitions of accuracy, including per-document accuracy, per-field accuracy, weighted-field accuracy, and priority-enforced field accuracy.