Skip to main content
Version: 10.2.8

Post-processing acceleration

Normalizer Search Engine tests different combinations of normalizers for each field to select the optimal chain. The final choice is the most useful and the safest for particular case. This minimizes the practice of designing custom Post-processors, so MLE doesn't have to write a Post-processor in the majority of the cases.

Normalizer Search Engine works especially well on numeric fields: numbers, dates, amount, and price. Quality of these fields can improve up to 70% in comparison with the same IE flow without the post-processing. Normalizer Search Engine is also useful for select one, email, currency, address, text fields etc.

Text normalization for all text fields

UpperCaseNormalizer

Modifies text to contain only upper-case letters for the specified locale.

  • Before: aBCdE
  • After: ABCDE

LowerCaseNormalizer

Modifies text to contain only lower-case letters for the specified locale.s

  • Before: aBCdE
  • After: abcde

CapitalizeNormalizer

Capitalizes all the whitespace separated words in a string. Only the first letter of each word is changed.

  • Before: aBC bcC cCD
  • After: ABC BcC CCD

TrimmerNormalizer

Removes spaces at the beginning and the end of the string.

  • Before: " aa a "
  • After: "aa a"

WhitespaceNormalizer

Replaces multiple occurrences of any whitespace symbol with one standard whitespace symbol.

  • Before: "cat dog"
  • After: "cat dog"

WhitespaceRemoveNormalizer

Replaces all occurrences of any whitespace symbol.

  • Before: "cat dog "
  • After: "catdog"

PunctuationRemoveNormalizer

Removes all occurrences of punctuation marks.

  • Before: cat,dog
  • After: catdog

MapDocumentTextToValueNormalizer

Text normalization for select one/optional one field types based on the automatic dictionary provided in resources and the nearest neighbors strategy. Substitutes the answer for the value from the dictionary, if the similarity score between them is acceptable.

Formatting

DateFormatNormalizer

Date format normalization for date fields based on automatic date-month order determination for given locale. If date format pattern is not provided, the default one ("MM/dd/yyyy") will be used.

  • Before: 1989,10,22
  • After: 10/22/1989

AmountFormatNormalizer

Amount formatting normalization for amount fields based on automatic locale determination used for delimeter choice. If amount format pattern is not provided, the default one ("#.00") will be used.

  • Before: 4,32
  • After: 4.32

DecimalNumberFormatNormalizer

Number formatting normalization for number fields based on automatic locale determination (used for delimiter choice).

  • Before: 3,14159265359
  • After: 3.14159265359

OCR error correction

OcrDateFixNormalizer

Removes OCR errors from the date string.

  • Before: Ju1 20, 2015
  • After: Jul 20, 2015

DictionaryCorrectionNormalizer

Removes OCR errors from the string based on the dictionary automatically gathered by GoldDataAnalyzers.

  • Before: l 5QO
  • After: 1 500

UsStatesAbbreviationNormalizer

Replaces the name of the US state by its abbreviation for address fields.

  • Before: Oklahoma
  • After: OK

For Data, Amount and Number fields, formatting pattern should be provided in the field info normalizationPattern attribute, otherwise the default pattern ("MM/dd/yyyy" for Date fields, "#.00" for Amount fields) will be used:

new FieldInfo.Builder("invoice_date")
.type(FieldType.DATE)
.multiValue(false)
.property(GenericPipelineConfiguration.PARAM_NORMALIZATION_PATTERN, "MM/dd/yyyy")
.build();

Configuration

The new components include:

  • GoldDataAnalyzersConfiguration: a set of analyzers grouped by field type. Gold Data Analyzer finds patterns and similarities in training data and stores them in resources.
  • GenericNormalizersConfiguration: a set of normalizers grouped by field type. Normalizers Search Engine uses it to find the best chain of normalizers. Normalizers use the output of gold data analyzers.
  • BestNormalizersProcessor: reads the serialized best normalizers chain and applies them to corresponding fields.
@ModelConfiguration()
@Import(configurations = {
@Import.Configuration(CandidateFeatureExtractorsConfiguration.class),
@Import.Configuration(GenericIeAnnotatorConfiguration.class),
@Import.Configuration(GenericNormalizersConfiguration.class),
@Import.Configuration(GoldDataAnalyzersConfiguration.class),
@Import.Configuration(GenericHpoConfiguration.class),
@Import.Configuration(GenericIePipelineConfiguration.class),
@Import.Configuration(GenericIeMlConfiguration.class),
@Import.Configuration(GenericIeParserConfiguration.class)
})
public class GenericIeHypermodelConfiguration {

@Named("basePostProcessors")
public List<Processor> getPostProcessors(DefaultConfigurationContext configurationContext) {
return Lists.newArrayList(new BestNormalizersProcessor(configurationContext.getResource("/lookup/normalizers")),
new CheckerModelPostProcessor(),
new DataValueNormalizationProcessor(),
new RowBasedGroupingProcessor());
}
}

Execution flow

  1. Gold Data Analyzer specified for the selected field is executed. The output files (local.json, ocr-corrections.csv) are stored in the resources directory.
  2. All possible combinations and permutations of fields normalizers are generated.
  3. Statistics is calculated and stored for all generated combinations in the normalizer-statistics.csv file.
  4. Best normalizers are serialized in the best-normalizers.json file.
  5. The best-normalizers.json file is used by BestNormalizersProcessor for post processing.

Here's the structure of files:

output
└──resources
└──lookup
└──normalizers
├──email
│ └──normalizer-statistics.csv
├──invoice_date
│ ├──best-normalizers.json
│ ├──date-format-priority.txt
│ └──normalizer-statistics.csv
├──invoice_number
│ ├──best-normalizers.json
│ ├──normalizer-statistics.csv
│ └──ocr-corrections.csv
├──price
│ ├──best-normalizers.json
│ ├──locale.json
│ └──normalizer-statistics.csv
├──product
│ └──normalizer-statistics.csv
├──quantity
│ ├──best-normalizers.json
│ ├──locale.json
│ └──normalizer-statistics.csv
└── total_amount
├──best-normalizers.json
├──locale.json
└──normalizer-statistics.csv