Gather ML requisites
The target audience for this course are resources that work directly with customers and are in a position to provide education on WorkFusion's technology and guidance on adopting Automation within their organization. The Primary Audience are resources that are typically engaged during pre-sales stage of the sales cycle with the objective of making the technical close and solutioning the initial project.
Task details
Before you begin with estimation, the following information is required:
Goal: What is the expectation at the end of this project?
Scope:
Collect business requirements, determine model types:
- Classification models (number of models, number of categories in each)
- Information Extraction (number of models, number of fields to extract in each)
Data Volume:
- Typical Monthly Volume
- Project Volume
Data formats: TXT, PDF, TIFF, XLS, DOC
- Is conversion for documents required?
- Is image quality good or pre-processing required?
- How many pages does each document contain in averedge?
Data output
- Field Name
- Description
- Data Type
- Mapped master Data
Types of Machine Learning tasks
- Data Extraction: Headers Data, Line Items
- Matching:
- Document/Email Classification
- Sub-Document Identification
- Data Mapping
- Taxonomy
Deployment:
- On-prem
- Cloud
- SPA Sandbox
Invoice processing example
- Data Volume
- Typical Monthly Volume: 5,000 documents
- Project Volume: 1,000 documents
- Data formats: PDF
- Data output: 5 fields:
- Invoice Number, Number, As Found on Invoice
- Invoice Date, European Style DD-MMM-YYYY, As found on Invoice but formatted
- Invoice Amount, Number, As found on Invoice
- Invoice Currency, GBP / USD / EUR / OTH, Mapped to 4 values
- Vendor Name, As found on Invoice
- Types of Machine Learning tasks: Data Extraction, Header Data only
- Deployment: On-Prem
Data example
For full assessment:
- Classification
- 20-30 example emails
- Classification List (Done)
- Information Extraction
- 20-30 example emails/attachments
- Field list