> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Workshop 7: Identity & Matching

> Define who the firm is, and decide when two mentions are the same client, person or project.

This workshop answers two questions Experio cannot answer on its own. First: **who is the firm?** If Experio does not know the firm's own names, it extracts the firm as just another external company in every proposal and contract. Second: **when are two records the same thing?** Documents mention "Lakeshore Health, Inc.", "Lakeshore" and "LSH". Resumes say "Bob Chen" where HR says "Robert Chen". Each entity type needs a rule for when mentions merge, when a person reviews them, and when a new node is created. Too loose, and two different people become one. Too strict, and one client becomes five. Either way, answers are wrong.

The outputs are the [client configuration](/implementation/key-concepts#client-configuration) and a [matching strategy](/implementation/key-concepts#matching-strategy) per entity type, plus an owner for the [conflict resolution](/implementation/key-concepts#conflict-resolution) queue.

## At a glance

| | |
| - | - |
| **Purpose** | Set the firm's identity; agree matching methods, normalization, filters and thresholds per entity type; staff the human review queue |
| **When** | Phase 2 (Model), weeks 3–4, after [Workshop 6: Data Mapping](/implementation/ws-data-mapping) |
| **Duration** | 2 hours |
| **Attendees** | Experio: FFE (leads), Experio SMEs as needed. Client: PM, super user, one SME per high-value entity type (for example, BD for clients, resource management for people), whoever will staff the review queue |
| **Inputs** | Saved ontology; agreed match keys from Workshop 6; **variant cards** (see homework); the firm's official name and known abbreviations |
| **Outputs** | Client Configuration entered; a matching strategy per entity type with a written rationale; review queue owners and SLA |
| **Admin pages** | **Administer > Client Configuration**, **Process > Matching Strategies**, **Process > Conflict Resolution** |

## Before the workshop

**FFE and Experio SMEs**

* List every entity type that documents set to **Create or Match**. Only these go through matching. Entity types that documents set to **Match Only**, and all structured loads, skip it, because structured matching is exact only.
* Check which entity types have a vector index (on the Ontology page, **Indexes** tab). Vector Similarity matching works only for those.
* Print or prepare the variant cards the client sends.

**Client homework: variant cards**

For each entity type that documents will match (at minimum Client and Employee), send **20 real pairs of names** taken from actual documents and systems. Each card holds two mentions and where each came from. Include:

* Legal suffixes: "Lakeshore Health" / "Lakeshore Health, Inc."
* Abbreviations and acronyms: "Meridian Bank" / "MB"
* Nicknames: "Robert Chen" / "Bob Chen"
* Typos and OCR errors: "Meridan Bank"
* Same name, different things: two different "J. Patel"s; "Lakeshore Health" and "Lakeshore Health Partners"
* Former names after a merger or rebrand

The hard cases matter more than the easy ones. Ask SMEs for the pairs that confuse *them*.

## Agenda

| Time | Item |
| - | - |
| 0:00 | Firm identity: canonical name, synonyms, display name, logo (15 min) |
| 0:15 | How matching works: methods, scores, thresholds, the review queue (15 min) |
| 0:30 | Variant card exercise, Client (25 min) |
| 0:55 | Variant card exercise, Employee (25 min) |
| 1:20 | Other entity types: exact key or match? (15 min) |
| 1:35 | Cost of errors per entity type; set thresholds (15 min) |
| 1:50 | Staff the review queue; agree SLA (10 min) |

## Running the workshop

### Part A: the firm's own identity

Ask the room:

* "What is the firm's official name, exactly as it should appear?"
* "What else do people call it in documents? Abbreviations, old names, the name on letterhead, the name in email signatures?"
* "What should users see in the Experio header?"

The canonical name and synonyms are given to the extraction step, so mentions of the firm are recognised as *us* and not extracted as an external company. Miss a common synonym and every proposal can produce a Client node for your own firm.

### Part B: entity resolution

<Steps>
  <Step title="Explain what matching does">
    When a document mentions an entity set to **Create or Match**, Experio looks for existing candidates using the enabled methods and scores each one from 0 to 1:

    | Method | Catches |
    | - | - |
    | **Exact** | Same value, case-insensitive |
    | **Vector Similarity** | Same meaning, different wording (needs a vector index on the entity type) |
    | **Fuzzy** | Typos, small spelling differences, abbreviations |
    | **Phonetic** | Names that sound alike ("Katherine" / "Catherine") |
    | **Synonym** | Values listed in the node's `synonyms` attribute |

    Each method has a weight. The score decides what happens next:

    | Score | Default | What happens |
    | - | - | - |
    | Auto-Match and above | 0.9 | Merged automatically |
    | LLM Disambiguation up to Auto-Match | 0.7–0.9 | An AI model compares the two and decides |
    | Human Review up to LLM Disambiguation | 0.5–0.7 | Queued in **Process > Conflict Resolution**; a provisional node is created meanwhile |
    | Below Human Review | below 0.5 | New node |
  </Step>

  <Step title="Run the variant cards">
    Deal the cards for one entity type. For each card the room decides: **same**, **different**, or **can't tell without more context**. Record the decision and the reason on the card.

    Then sort the cards into piles by *why* they differ: suffix, abbreviation, nickname, typo, genuinely different. Each pile points to a setting:

    * Suffix pile → **Remove Company Suffixes** normalization.
    * Abbreviation pile → **Expand Abbreviations**, or put the abbreviation in the node's `synonyms` attribute and enable **Synonym**.
    * Nickname and typo piles → **Fuzzy** and **Phonetic**.
    * "Can't tell" pile → these belong in the human review band. Ask: "What extra fact would settle it?" If the answer is "which practice they're in" or "the date", that is a filter.
    * "Different" pile with similar names → these are why the auto-merge threshold must stay high.
  </Step>

  <Step title="Add filters where context decides">
    * **Relationship filter**: only consider candidates connected to a given parent type or through given relationships. Example: an Employee mention is compared with employees in the same Practice.
    * **Temporal filter**: penalise candidates whose dates are far apart (**Max Distance (Days)**, **Penalty Per Year**, **Date Attribute Names**). Useful for entities that recur with the same name over years, such as annual proposals.
  </Step>

  <Step title="Weigh the cost of each kind of error">
    Ask for each entity type: "Which is worse here, a false merge or a duplicate?"

    | Entity type | False merge (two things become one) | Duplicate (one thing becomes two) | Lean |
    | - | - | - | - |
    | Employee | Severe: one person's projects and skills show up on another person's profile; staffing answers go to the wrong person | Moderate: a person's history is split, so answers are incomplete | Strict: higher auto-merge, wide human band |
    | Client | Severe for contracts and conflicts of interest | Moderate: past performance is split across two nodes | Balanced, with strong normalization and synonyms |
    | Project | Severe: finance figures attach to the wrong engagement | Avoid entirely with an exact key | Exact key, documents set to **Match Only** |
    | Skill | Low | Low (enrichment and taxonomies absorb it) | Loose |

    A false merge is usually worse than a duplicate. A duplicate is incomplete but true. A false merge is confidently wrong, and it is harder to find and undo.
  </Step>

  <Step title="Set thresholds from the cards">
    Score the cards mentally against the proposed settings. Every "different" card must fall below Auto-Match. Every clear "same" card should land at or above LLM Disambiguation. "Can't tell" cards should land in the human band. Adjust thresholds until the piles fall where the room put them. Tune again after the [pilot ingestion](/implementation/pilot-ingestion) with real scores.
  </Step>

  <Step title="Staff the review queue">
    Ask: "Who knows these names well enough to decide, and how quickly can they do it?" For each entity type, name a primary and a backup reviewer and agree an SLA. Reviewers open **Process > Conflict Resolution** and choose **Merge**, **Confirm New** or **Skip**. Until someone decides, the provisional node stays in the graph, so a slow queue means duplicates in answers. Experio does not send reminders or enforce the SLA. The PM checks queue size in the weekly status meeting.
  </Step>
</Steps>

## Worked example: Northbridge Consulting

**Client Configuration**: Organization Name "Northbridge Consulting"; Synonyms "Northbridge", "NBC", "Northbridge Consulting Group". Marcus added the display name and logo.

**Matching strategies**

| Strategy | Methods | Normalization | Filters | Thresholds | Why |
| - | - | - | - | - | - |
| Default | Exact + Vector Similarity | Unicode, punctuation | — | 0.9 / 0.7 / 0.5 | Safe baseline for types not listed below |
| Client | Exact + Fuzzy + Synonym | Remove Company Suffixes, Remove Prefixes, punctuation | — | 0.9 / 0.7 / 0.5 | "Lakeshore Health, Inc." = "Lakeshore Health". Sales ops adds CRM aliases ("LSH") to the Client `synonyms` attribute through `accounts.csv` |
| Employee | Exact + Fuzzy + Phonetic | Unicode, punctuation | Relationship filter to Practice (optional) | 0.9 / 0.7 / 0.5, human band kept | Resumes say "Bob Chen", HR says "Robert Chen". Two "J. Patel"s exist, so no lowering of Auto-Match |
| Project | Exact on `project_code` | — | — | — | Documents set Project to **Match Only**, so a status report can never create a project Finance doesn't know |

To make **Synonym** work for Client, the team added a `synonyms` list attribute to Client in the ontology and mapped it from an `aliases` column in `accounts.csv`. The Synonym method compares a mention with each node's `name` and its `synonyms` list.

**Variant card results (Employee, 20 cards)**: 11 same, 5 different, 4 can't tell. The "different" pile included "J. Patel" (Strategy) and "J. Patel" (Technology). This is why the room kept the Practice relationship filter in mind for the pilot.

**Review queue**: Marcus Lee is primary for all types. Sam Whitfield is backup for Employee, Dana Ortiz for Client. SLA: two business days during the pilot, one business day during full ingestion.

## Accelerate with AI

There is no AI helper that sets matching strategies. The LLM Disambiguation band is itself AI help at run time. **Admin Copilot** (⌘J) can explain each setting from these docs while you configure it.

## Entering it in Experio

<Steps>
  <Step title="Firm identity">
    **Administer > Client Configuration**: set Organization Name and Synonyms, then Display Name and Logo (PNG or JPG). See [Client Configuration](/admin-guide/client-configuration).
  </Step>

  <Step title="Matching strategies">
    **Process > Matching Strategies**: edit the default strategy, then **Create New Strategy** for each entity type that needs its own rules. Set methods and weights, filters, normalization and thresholds. Each section has **Reset** if you need to start again. See [Matching Strategies](/admin-guide/matching-strategies).
  </Step>

  <Step title="Processing type on artifact types">
    On each artifact type, confirm each entity's processing type: **Create or Match** where matching should run, **Match Only** where documents must never create the entity.
  </Step>

  <Step title="Review queue">
    Give reviewers access to **Process > Conflict Resolution** (Process Managers group). See [Conflict Resolution](/admin-guide/conflict-resolution) and [Roles & Permissions](/implementation/roles-and-permissions).
  </Step>
</Steps>

## Common pitfalls

* **Forgetting the firm's own synonyms.** "NBC" missing from Client Configuration gives a Client node called "NBC" linked to every proposal.
* **Lowering Auto-Match to clear the queue.** This trades a visible queue for invisible false merges.
* **Expecting matching to fix structured data.** Structured loads are exact only. Fix formats in the export ([Workshop 6](/implementation/ws-data-mapping)).
* **Enabling Vector Similarity on a type with no vector index.** It finds nothing. Add the index on the Ontology page first.
* **Synonym method with no `synonyms` values.** The method reads the node's `synonyms` attribute. If nothing fills it, it adds nothing.
* **An unstaffed queue.** Provisional nodes pile up and duplicates show in answers.
* **Using only easy variant cards.** The room agrees quickly and the thresholds are never tested on the cases that matter.

## Exit criteria

* [ ] Organization Name and all known synonyms are entered in Client Configuration.
* [ ] Variant cards are decided for each entity type set to Create or Match, and kept with the project records.
* [ ] Each such entity type has a matching strategy (or uses the default on purpose), with a one-line rationale.
* [ ] Entity types with exact keys are set to Match Only on artifact types where appropriate.
* [ ] Vector indexes exist for every type that uses Vector Similarity.
* [ ] Review queue owners, backups and SLA are named and agreed.

## Next

Continue to [Workshop 8: Enrichment Rules](/implementation/ws-enrichment-rules), which runs in Phase 3 once the pilot ingestion has put real data in the graph.
