Back to all guides
Entity Matching

Entity Matching: Link Records That Describe the Same Entity

How to configure and run Entity Matching Pipelines that connect nodes describing the same real-world entity with a _SAME_AS relationship.

What is Entity Matching?

Most Identity Knowledge Graphs (IKG) are fed from several independent sources: a customer database, an HR system, a partner feed. Each source ingests its own node type, and nothing links a Customer to the Employee who is the same person. Entity Matching finds those pairs. It compares the property values of two node types, computes a similarity score per pair, and connects every pair above a threshold with a _SAME_AS relationship.

Two cases are covered:

  • Cross-type linking: nodes of different types that refer to the same entity, e.g. Customer and Employee.
  • Duplicate detection: nodes of the same type that represent the same entity, e.g. two Employee records.

The result is a unified view of the graph without merging or rewriting the source nodes: the original records stay as they are, and the _SAME_AS edge carries the evidence.

How does it work?

An Entity Matching Pipeline is a configuration object that names a source node type, a target node type, and a similarity threshold. It runs in two steps:

  1. Property mapping starts automatically when the pipeline is created. The platform compares the property names of the two types and proposes which source property should be compared with which target property, each pair with a score for how similar the names are. This step is asynchronous; its state is property_mapping_status.
  2. Matching is started explicitly through the Entity Matching REST API. Using the stored mapping, or a custom one sent with the run, it scores every source/target node pair and writes a _SAME_AS relationship for each pair whose score is above the cutoff. Its state is entity_matching_status.

Both statuses take the values PENDING, IN_PROGRESS, SUCCESS, or ERROR. When a step ends in ERROR, the matching property_mapping_message or entity_matching_message on the pipeline explains why.

There is no scheduler. The rerun_interval field exists on the configuration but is not supported; every run is triggered by a call.

What credentials do I need?

  • Creating, reading, updating, deleting pipelines: Service Account credentials (Config API, Authorization: Bearer).
  • Reading the suggested mapping, running, and polling status: AppAgent credentials (X-IK-ClientKey) with the EntityMatching API permission. Without it, every call under /entity-matching/v1/ fails with 401 and insufficient API access level for appAgent.
  • Ingesting the nodes to match: AppAgent credentials (Capture API).

Configuration methods: REST Config API (Config API documentation) and the Hub.

Pipeline Configuration (Config API)

REST API Endpoints

Operation Method Endpoint
Create POST /configs/v1/entity-matching-pipelines
Read by ID GET /configs/v1/entity-matching-pipelines/{id}
Read by name GET /configs/v1/entity-matching-pipelines/{name}?location={project_id}
List GET /configs/v1/entity-matching-pipelines?project_id={id} (optional full_fetch=true, search=)
Update PUT /configs/v1/entity-matching-pipelines/{id} (If-Match: ETag)
Delete DELETE /configs/v1/entity-matching-pipelines/{id} (If-Match: ETag)

Create Request Syntax

{
  "project_id": "<string>",
  "name": "<string>",
  "display_name": "<string>",
  "description": "<string>",
  "node_filter": {
    "source_node_types": [
      "<string>"
    ],
    "target_node_types": [
      "<string>"
    ]
  },
  "similarity_score_cutoff": <number>
}

What does each field mean?

Required fields

  • project_id: The GID of the project that owns the pipeline and the IKG it scans.
  • name: Unique, immutable identifier. It is also written to every _SAME_AS relationship the pipeline creates (see below).
  • node_filter.source_node_types: Node types to match. At least one, unique. Only the first entry is used by the current implementation.
  • node_filter.target_node_types: Node types to match against. Same rules. Use the same type on both sides to detect duplicates.
  • similarity_score_cutoff: Number between 0 and 1. Pairs whose similarity score is above it are connected.

Optional fields

  • display_name: 2-254 characters. Equals name when not set.
  • description: 2-65000 characters.
  • rerun_interval: Accepted but not supported; it does not schedule anything.

Update semantics. PUT accepts similarity_score_cutoff, display_name, description, and rerun_interval. node_filter is immutable: to match different types, create a new pipeline. For display_name and description, omitted or null keeps the current value and an empty string removes it. A stale If-Match value returns 412.

Example: Create Entity Matching Pipeline

POST /configs/v1/entity-matching-pipelines
{
  "project_id": "gid:AAAABbbbCCCC...",
  "name": "employee-customer-match",
  "display_name": "Employees who are also customers",
  "description": "Links Employee nodes to the Customer node describing the same person",
  "node_filter": {
    "source_node_types": [
      "Employee"
    ],
    "target_node_types": [
      "Customer"
    ]
  },
  "similarity_score_cutoff": 0.7
}

Read Response

Reading a pipeline returns the configuration plus the state of both steps and the outcome of the last run:

{
  "id": "gid:AAAABbbbCCCC...",
  "name": "employee-customer-match",
  "display_name": "Employees who are also customers",
  "description": "Links Employee nodes to the Customer node describing the same person",
  "node_filter": {
    "source_node_types": [
      "Employee"
    ],
    "target_node_types": [
      "Customer"
    ]
  },
  "similarity_score_cutoff": 0.7,
  "property_mapping_status": "SUCCESS",
  "property_mapping_message": "",
  "property_mappings": [
    {
      "source_node_type": "Employee",
      "source_node_property": "email",
      "target_node_type": "Customer",
      "target_node_property": "email",
      "similarity_score_cutoff": 1
    },
    {
      "source_node_type": "Employee",
      "source_node_property": "family_name",
      "target_node_type": "Customer",
      "target_node_property": "family_name",
      "similarity_score_cutoff": 1
    }
  ],
  "entity_matching_status": "SUCCESS",
  "entity_matching_message": "",
  "last_run_time": "2026-09-28T09:12:00Z",
  "matched_entities": 2,
  "report_url": "...",
  "report_type": "...",
  "rerun_interval": "",
  "organization_id": "gid:...",
  "project_id": "gid:...",
  "create_time": "2026-09-28T09:00:00Z",
  "update_time": "2026-09-28T09:12:00Z",
  "created_by": "gid:...",
  "updated_by": "gid:..."
}
  • property_mappings[]: the rules the next run will use unless overridden. The per-pair similarity_score_cutoff is how similar the two property names are (1 means identical), not a data similarity. Use it to decide which suggested pairs to keep.
  • matched_entities: number of _SAME_AS relationships created by the last run.
  • last_run_time: when the last run was started. report_url and report_type point at the analysis report of that run.
  • On the list endpoint, pass full_fetch=true to get these objects instead of metadata only.

Entity Matching REST API

Base URL: https://eu.api.indykite.com/entity-matching/v1 or https://us.api.indykite.com/entity-matching/v1. All three endpoints take the pipeline GID in the path and authenticate with X-IK-ClientKey. A pipeline that does not exist, or that belongs to another project than the agent's, answers 404 with entity matching pipeline was not found.

Step 1: Read the suggested property mapping

GET /entity-matching/v1/pipelines/{id}/property-mappings

Returns the mapping produced by the property mapping step, in the shape you can send back on a run:

{
  "id": "gid:AAAABbbbCCCC...",
  "suggested_property_mappings": [
    {
      "source_node_type": "Employee",
      "source_node_property": "email",
      "target_node_type": "Customer",
      "target_node_property": "email",
      "similarity_score_cutoff": 1
    }
  ]
}
Status Body Meaning
200 suggested_property_mappings Mapping is ready (property_mapping_status is SUCCESS).
202 {"message": "the property mapping process has not yet been started"} Status is PENDING. Poll again.
202 {"message": "the property mapping process is not yet completed"} Status is IN_PROGRESS. Poll again.
409 {"message": "the property mapping process failed", "errors": ["..."]} Status is ERROR; errors carries the property_mapping_message. Fix the data or recreate the pipeline.

Step 2: Run the pipeline

POST /entity-matching/v1/pipelines/{id}/runs
{
  "similarity_score_cutoff": 0.7,
  "custom_property_mappings": [
    {
      "source_node_property": "email",
      "target_node_property": "email"
    },
    {
      "source_node_property": "family_name",
      "target_node_property": "family_name"
    },
    {
      "source_node_property": "given_name",
      "target_node_property": "given_name"
    }
  ]
}
  • similarity_score_cutoff (required, 0 to 1): the threshold for this run. It is also stored on the pipeline.
  • custom_property_mappings (optional): property pairs to compare, each with source_node_property and target_node_property. The node types are taken from the pipeline's node_filter. When present, the list replaces the stored property_mappings; when absent, the stored mapping is used.

Response 202 Accepted:

{
  "id": "gid:AAAABbbbCCCC...",
  "last_run_time": "2026-09-28T09:12:00Z",
  "etag": "..."
}

Starting a run sets entity_matching_status to IN_PROGRESS and removes the _SAME_AS relationships the same pipeline created earlier. Each run therefore reflects the current data: a pair that no longer scores above the cutoff loses its relationship, and a pair that now does gains one. Relationships created by other pipelines are untouched.

Status Meaning
422 cannot run pipeline without property mapping, either wait for previous step to complete or pass a valid custom property mapping The property mapping step has not produced a mapping yet and no custom_property_mappings were sent.
422 with errors[] Body validation: similarity_score_cutoff missing or outside 0-1, a mapping entry missing a property name, or a path id that is not an Entity Matching Pipeline GID.
404 entity matching pipeline was not found Unknown ID, or a pipeline of another project.
412 The pipeline was modified concurrently; read it again and retry.

Step 3: Poll the status

GET /entity-matching/v1/pipelines/{id}/status
{
  "id": "gid:AAAABbbbCCCC...",
  "property_mapping_status": "SUCCESS",
  "entity_matching_status": "IN_PROGRESS"
}

Poll until entity_matching_status is SUCCESS or ERROR. The error text is not part of this response; read the pipeline through the Config API for entity_matching_message and matched_entities.

Results in the IKG

A successful run writes one _SAME_AS relationship per matched pair, from the source node to the target node:

(Employee:employeeA)-[:_SAME_AS]->(Customer:customerA)

The relationship carries these properties:

Property Meaning
_final_score The similarity score of the pair, 0 to 1. Always above the cutoff of the run that created it.
_ingress The name of the pipeline that created it.
pipeline_id The pipeline's ID.
_service entity-matching.

Rules that follow from the relationship being platform-managed:

  • Capture cannot create or delete it. _SAME_AS does not match the relationship type pattern ^[A-Z]+(?:_[A-Z]+)*$, so POST /capture/v1/relationships and the relationship delete endpoints reject it. Runs and pipeline deletion are the only ways to change it.
  • Deleting the pipeline deletes its relationships. DELETE /configs/v1/entity-matching-pipelines/{id} removes every _SAME_AS relationship carrying that pipeline_id.
  • It is not part of the Data Schema. GET /data-schema/v1/ skips platform-generated relationship types, so _SAME_AS does not appear there even when matches exist.
  • Several pipelines can link the same nodes. Each writes its own relationship, distinguished by _ingress and pipeline_id.

Reading matches with ContX IQ

The relationship is traversable in a ContX IQ policy like any other. This policy lets an application list, for an employee, the customer records that describe the same person and the score of each match:

{
  "meta": {
    "policy_version": "1.0-ciq"
  },
  "subject": {
    "type": "_Application"
  },
  "condition": {
    "cypher": "MATCH (subject:_Application) MATCH (employee:Employee)-[same:_SAME_AS]->(customer:Customer)",
    "filter": {
      "attribute": "employee.external_id",
      "operator": "=",
      "value": "$employee_external_id"
    }
  },
  "allowed_reads": {
    "nodes": [
      "employee.external_id",
      "customer.external_id",
      "customer.property.email"
    ],
    "relationships": [
      "same"
    ]
  }
}

Knowledge Query:

{
  "nodes": [
    "employee.external_id",
    "customer.external_id",
    "customer.property.email"
  ],
  "relationships": [
    "same"
  ]
}

The returned same relationship exposes _final_score, _ingress, and pipeline_id in its Props. Filter on _ingress when more than one pipeline links the same types.

Operational notes

  • The suggested mapping is derived from property names. Sources that spell the same attribute differently (surname vs family_name) will not be paired automatically; send the pair in custom_property_mappings.
  • The mapping is computed once, at creation. When the data model changes (new property types, renamed properties), create a new pipeline so the mapping reflects the current schema.
  • Run after ingestion has finished, not during it. A run scores the nodes present at that moment and removes the pipeline's earlier relationships, so a half-loaded graph produces an incomplete result until the next run.
  • A run is asynchronous and scales with the number of source and target nodes. Poll the status endpoint rather than assuming completion.
  • Lower similarity_score_cutoff means more matches and more false positives. Start from the suggested pairs with a name similarity of 1 and a cutoff around 0.7, inspect _final_score on the results, then adjust.

Error Handling

Config API

HTTP Code Meaning Common Cause
400 Bad Request Invalid JSON, missing required fields
401 Unauthorized Invalid or missing Bearer token
403 Forbidden Insufficient permissions for the project
404 Not Found Pipeline ID/name doesn't exist
412 Precondition Failed ETag mismatch (concurrent modification)
422 Unprocessable Entity Validation error: empty node_filter lists, duplicate types, cutoff outside 0-1, attempt to change node_filter

Entity Matching REST API

HTTP Code Meaning
401 insufficient API access level for appAgent The agent lacks the EntityMatching permission.
401 missing app agent credentials No X-IK-ClientKey header.
404 entity matching pipeline was not found Unknown ID, or a pipeline of another project.
202 Run accepted, or property mapping still pending / in progress (see the per-endpoint tables above).
409 the property mapping process failed Mapping step ended in ERROR; details in errors[].
422 Run body invalid, or no mapping available and none supplied.

Next Steps