Skip to content

Repository files navigation

SPIDER: Synthetic Person Information Dataset for Entity Resolution

This repository serves as a public landing page and documentation hub for the SPIDER dataset.

SPIDER stands for Synthetic Person Information Dataset for Entity Resolution. It is a synthetic person information dataset designed for entity resolution, duplicate detection, record linkage, and data-quality testing research.

Note: This repository does not host the dataset files. The official dataset is archived and versioned on Figshare using the DOI listed below.

Official Dataset DOI

The official Version 2 dataset is archived on Figshare:

Version 2 DOI: https://doi.org/10.6084/m9.figshare.29595599.v2

Associated Published Paper

The dataset is described in the following published IEEE Data Description article:

Chinnappa, P., Mathur, Y., & Arokiyadass, R. Descriptor: Synthetic Person Information Dataset for Entity Resolution (SPIDER). IEEE Data Descriptions, 2026.

Paper DOI: https://doi.org/10.1109/IEEEDATA.2026.3663349

Purpose of This Repository

This GitHub repository is intended to improve discoverability and provide supporting documentation for the SPIDER dataset.

It includes:

  • Dataset overview
  • Duplicate-generation rule summary
  • Data dictionary
  • Validation summary
  • Citation instructions
  • Link to the official Figshare dataset archive
  • Reference to the associated published IEEE Data Description article

Dataset Summary

SPIDER contains 50,000 synthetic person records, including:

  • 40,000 base records
  • 10,000 controlled duplicate records

The dataset was created to support research and testing in:

  • Entity resolution
  • Duplicate detection
  • Record linkage
  • Synthetic test data generation
  • Data quality validation
  • Software testing and QA research

Duplicate Generation Rules

Controlled duplicates were generated using seven rule categories:

  1. Email variation
  2. Phone number variation
  3. Last name variation
  4. Address variation
  5. Nickname variation
  6. First name spelling variation
  7. Shared account variation

These rule categories were designed to represent common duplicate patterns found in person-information systems, CRM platforms, QA test environments, and entity resolution workflows.

Accessing the Dataset

The full dataset is available through Figshare:

Version 2 DOI: https://doi.org/10.6084/m9.figshare.29595599.v2

Please use the Figshare Version 2 DOI when accessing, referencing, or citing the dataset.

Important Note About Data Hosting

This repository does not contain CSV files, sample records, or any copy of the dataset.

The purpose of this repository is to serve as a documentation and discoverability page. Figshare remains the official archived source for the dataset.

Citation

If you use the SPIDER dataset, please cite both the associated IEEE Data Description article and the official Figshare Version 2 dataset DOI.

Associated Paper

Chinnappa, P., Mathur, Y., & Arokiyadass, R. Descriptor: Synthetic Person Information Dataset for Entity Resolution (SPIDER). IEEE Data Descriptions, 2026.

Paper DOI: https://doi.org/10.1109/IEEEDATA.2026.3663349

Dataset DOI

Chinnappa, P., Mathur, Y., & Arokiyadass, R. SPIDER: Synthetic Person Information Dataset for Entity Resolution. Figshare. Version 2.

Dataset DOI: https://doi.org/10.6084/m9.figshare.29595599.v2

Suggested Use Cases

SPIDER may be useful for:

  • Evaluating entity resolution algorithms
  • Testing duplicate detection logic
  • Validating record linkage workflows
  • Demonstrating synthetic test data generation
  • Supporting software QA and data-quality research
  • Benchmarking matching strategies using controlled duplicate clusters

Keywords

Synthetic data, entity resolution, duplicate detection, record linkage, data quality, test data generation, Faker API, software testing, QA research.

Repository Scope

This repository is limited to public documentation and discovery support for the SPIDER dataset.

The official dataset files, versioning, and archival record are maintained on Figshare.

About

Public landing page and documentation for the SPIDER synthetic person information dataset for entity resolution.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors