Skip to main content
This workshop answers two questions Experio cannot answer on its own. First: who is the firm? If Experio does not know the firm’s own names, it extracts the firm as just another external company in every proposal and contract. Second: when are two records the same thing? Documents mention “Lakeshore Health, Inc.”, “Lakeshore” and “LSH”. Resumes say “Bob Chen” where HR says “Robert Chen”. Each entity type needs a rule for when mentions merge, when a person reviews them, and when a new node is created. Too loose, and two different people become one. Too strict, and one client becomes five. Either way, answers are wrong. The outputs are the client configuration and a matching strategy per entity type, plus an owner for the conflict resolution queue.

At a glance

Before the workshop

FFE and Experio SMEs
  • List every entity type that documents set to Create or Match. Only these go through matching. Entity types that documents set to Match Only, and all structured loads, skip it, because structured matching is exact only.
  • Check which entity types have a vector index (on the Ontology page, Indexes tab). Vector Similarity matching works only for those.
  • Print or prepare the variant cards the client sends.
Client homework: variant cards For each entity type that documents will match (at minimum Client and Employee), send 20 real pairs of names taken from actual documents and systems. Each card holds two mentions and where each came from. Include:
  • Legal suffixes: “Lakeshore Health” / “Lakeshore Health, Inc.”
  • Abbreviations and acronyms: “Meridian Bank” / “MB”
  • Nicknames: “Robert Chen” / “Bob Chen”
  • Typos and OCR errors: “Meridan Bank”
  • Same name, different things: two different “J. Patel”s; “Lakeshore Health” and “Lakeshore Health Partners”
  • Former names after a merger or rebrand
The hard cases matter more than the easy ones. Ask SMEs for the pairs that confuse them.

Agenda

Running the workshop

Part A: the firm’s own identity

Ask the room:
  • “What is the firm’s official name, exactly as it should appear?”
  • “What else do people call it in documents? Abbreviations, old names, the name on letterhead, the name in email signatures?”
  • “What should users see in the Experio header?”
The canonical name and synonyms are given to the extraction step, so mentions of the firm are recognised as us and not extracted as an external company. Miss a common synonym and every proposal can produce a Client node for your own firm.

Part B: entity resolution

1

Explain what matching does

When a document mentions an entity set to Create or Match, Experio looks for existing candidates using the enabled methods and scores each one from 0 to 1:Each method has a weight. The score decides what happens next:
2

Run the variant cards

Deal the cards for one entity type. For each card the room decides: same, different, or can’t tell without more context. Record the decision and the reason on the card.Then sort the cards into piles by why they differ: suffix, abbreviation, nickname, typo, genuinely different. Each pile points to a setting:
  • Suffix pile → Remove Company Suffixes normalization.
  • Abbreviation pile → Expand Abbreviations, or put the abbreviation in the node’s synonyms attribute and enable Synonym.
  • Nickname and typo piles → Fuzzy and Phonetic.
  • “Can’t tell” pile → these belong in the human review band. Ask: “What extra fact would settle it?” If the answer is “which practice they’re in” or “the date”, that is a filter.
  • “Different” pile with similar names → these are why the auto-merge threshold must stay high.
3

Add filters where context decides

  • Relationship filter: only consider candidates connected to a given parent type or through given relationships. Example: an Employee mention is compared with employees in the same Practice.
  • Temporal filter: penalise candidates whose dates are far apart (Max Distance (Days), Penalty Per Year, Date Attribute Names). Useful for entities that recur with the same name over years, such as annual proposals.
4

Weigh the cost of each kind of error

Ask for each entity type: “Which is worse here, a false merge or a duplicate?”A false merge is usually worse than a duplicate. A duplicate is incomplete but true. A false merge is confidently wrong, and it is harder to find and undo.
5

Set thresholds from the cards

Score the cards mentally against the proposed settings. Every “different” card must fall below Auto-Match. Every clear “same” card should land at or above LLM Disambiguation. “Can’t tell” cards should land in the human band. Adjust thresholds until the piles fall where the room put them. Tune again after the pilot ingestion with real scores.
6

Staff the review queue

Ask: “Who knows these names well enough to decide, and how quickly can they do it?” For each entity type, name a primary and a backup reviewer and agree an SLA. Reviewers open Process > Conflict Resolution and choose Merge, Confirm New or Skip. Until someone decides, the provisional node stays in the graph, so a slow queue means duplicates in answers. Experio does not send reminders or enforce the SLA. The PM checks queue size in the weekly status meeting.

Worked example: Northbridge Consulting

Client Configuration: Organization Name “Northbridge Consulting”; Synonyms “Northbridge”, “NBC”, “Northbridge Consulting Group”. Marcus added the display name and logo. Matching strategies To make Synonym work for Client, the team added a synonyms list attribute to Client in the ontology and mapped it from an aliases column in accounts.csv. The Synonym method compares a mention with each node’s name and its synonyms list. Variant card results (Employee, 20 cards): 11 same, 5 different, 4 can’t tell. The “different” pile included “J. Patel” (Strategy) and “J. Patel” (Technology). This is why the room kept the Practice relationship filter in mind for the pilot. Review queue: Marcus Lee is primary for all types. Sam Whitfield is backup for Employee, Dana Ortiz for Client. SLA: two business days during the pilot, one business day during full ingestion.

Accelerate with AI

There is no AI helper that sets matching strategies. The LLM Disambiguation band is itself AI help at run time. Admin Copilot (⌘J) can explain each setting from these docs while you configure it.

Entering it in Experio

1

Firm identity

Administer > Client Configuration: set Organization Name and Synonyms, then Display Name and Logo (PNG or JPG). See Client Configuration.
2

Matching strategies

Process > Matching Strategies: edit the default strategy, then Create New Strategy for each entity type that needs its own rules. Set methods and weights, filters, normalization and thresholds. Each section has Reset if you need to start again. See Matching Strategies.
3

Processing type on artifact types

On each artifact type, confirm each entity’s processing type: Create or Match where matching should run, Match Only where documents must never create the entity.
4

Review queue

Give reviewers access to Process > Conflict Resolution (Process Managers group). See Conflict Resolution and Roles & Permissions.

Common pitfalls

  • Forgetting the firm’s own synonyms. “NBC” missing from Client Configuration gives a Client node called “NBC” linked to every proposal.
  • Lowering Auto-Match to clear the queue. This trades a visible queue for invisible false merges.
  • Expecting matching to fix structured data. Structured loads are exact only. Fix formats in the export (Workshop 6).
  • Enabling Vector Similarity on a type with no vector index. It finds nothing. Add the index on the Ontology page first.
  • Synonym method with no synonyms values. The method reads the node’s synonyms attribute. If nothing fills it, it adds nothing.
  • An unstaffed queue. Provisional nodes pile up and duplicates show in answers.
  • Using only easy variant cards. The room agrees quickly and the thresholds are never tested on the cases that matter.

Exit criteria

  • Organization Name and all known synonyms are entered in Client Configuration.
  • Variant cards are decided for each entity type set to Create or Match, and kept with the project records.
  • Each such entity type has a matching strategy (or uses the default on purpose), with a one-line rationale.
  • Entity types with exact keys are set to Match Only on artifact types where appropriate.
  • Vector indexes exist for every type that uses Vector Similarity.
  • Review queue owners, backups and SLA are named and agreed.

Next

Continue to Workshop 8: Enrichment Rules, which runs in Phase 3 once the pilot ingestion has put real data in the graph.