Skip to main content
Version: 10.3.2

Delete duplicates

overview

Task: identify and delete duplicate documents in a dataset.

Jupiter Notebook: deduplication

Input: a local CSV file with links to documents

  1. Import the Deduplicator.

  2. Upload data.

  3. Launch the Deduplicator.

  4. Provide required data as indicated below and click Deduplicate.

    • Select a dataframe: specify the name of the dataframe that contains the dataset.

    • Hash column name: add the name of the column in the dataframe that will have a hash generated

    • Is duplicate: specify the name of the new column that will contain the result (True or False) for the duplicate check.

    • Document field name: enter the name of the column that contains the input for checking (the column with the links to the documents or tagged text).

In the output, you’ll see a column containing the True result for all duplicate documents.

You can filter out duplicate values and save the result to a file.