ML project questionnaire template
Process
Process description or service ticket link? * |
|
|---|---|
| How sensitive is the data? * | |
| Cloud vs On-prem? | |
| Remote vs Onsite work? | |
| Possible implementation Partner? | |
| What is the target level of quality for automation (accuracy, automation rate)? Can't be 100% | |
| POC / PROD / Prod pilot? |
System information
| Number of ML servers with CPU/RAM/HDD * | |
|---|---|
| Operation system (distributive and version)* | |
| WorkFusion version |
Data type and format
| What is the document format? (PDF, HTML, Excel) * | |
|---|---|
| What % of documents contains (1) plain text; (2) tables; (3) other formats? | |
| What language(s) are the input documents and distribution% ? * | |
| How many different types/templates of documents? * | |
| What is the average size of documents? (number of pages for pdf; number of rows for excel) |
Fields
| Does the field appear once or multiple times per document? | |
|---|---|
| How often does a field appear in documents? (e.g. a field appears in 25% of documents) | |
| In case of large documents - are the fields clustered close to each other or spread over multiple pages? |
OCR
| Are the documents directly searchable? * | |
|---|---|
| What % of the input documents require OCR? * | |
| What % of documents contain handwritting? How much handwriting per document? * | |
| What is the DPI of the image-based documents? |
WorkFusion automation preparation
| List of the fields to extract: * | |
|---|---|
| List of fields with reference data, dictionaries? * |