> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Key Concepts

> Definitions of the building blocks of an Experio implementation: what each one is, what it's used for, and where it lives

This page defines the terms used throughout the implementation. Each entry says **what it is**, **what it's used for**, gives an **example** from [Northbridge Consulting](/implementation/worked-example), and names **where it lives** in the admin.

## How the pieces fit together

```mermaid theme={null}
flowchart TD
    Q["Golden questions"] -->|"drive"| O["Ontology<br/>entity types, attributes, relationships"]
    T["Taxonomies<br/>controlled lists"] -->|"constrain values in"| O
    O -->|"subset per document family"| AT["Artifact types<br/>+ extraction policy"]
    O -->|"target of"| DM["Data mappings"]
    O -->|"per entity type"| MS["Matching strategies"]
    O -->|"read and write"| ER["Enrichment rules"]
    CC["Client configuration"] -->|"firm identity"| AT
    AT --> G[("Knowledge graph<br/>+ lineage")]
    DM --> G
    MS --> G
    ER --> G
    G --> A["Assistants, templates,<br/>agent flows"]
    O -.->|"every save re-checks"| C["Ontology compatibility"]
```

## Modeling concepts

### Knowledge graph

**What it is.** A database of **nodes** (things) connected by **relationships** (links). Each node has a type and attributes.

**What it's for.** Answering questions that cross sources and follow connections. For example: "who worked on healthcare projects for clients we have an MSA with". Keyword search can't answer that.

**Example.** `(Employee: Maria Gomez) -[WORKS_ON_PROJECT {role: "Engagement Lead"}]-> (Project: Lakeshore Cloud Migration) -[FOR_CLIENT]-> (Client: Lakeshore Health)`

### Ontology

**What it is.** The schema of the knowledge graph. It is the agreed list of the kinds of things the business cares about ([entity types](#entity-type)), the facts recorded about each ([attributes](#attribute)), and the links allowed between them ([relationships](#relationship)).

**What it's for.** Everything else is defined against it. Extraction only looks for what the ontology names. Data mappings only write to what it defines. Chat uses it to plan graph queries. A clear, question-driven ontology is the biggest single factor in answer accuracy.

**Example.** Northbridge's pilot ontology has eight entity types: Client, Project, Employee, Practice, Skill, Proposal, Contract and Obligation. It has thirteen relationships between them.

**Where.** **Model & Define > Ontology**. Every save creates a revision (see Revision History). Built in [Workshop 3](/implementation/ws-ontology).

### Entity type

**What it is.** A kind of thing in the ontology, and the label on its nodes in the graph. It is written in singular PascalCase: `Client`, `Project`, `Employee`.

**What it's for.** It groups records that share attributes, relationships, a matching strategy and access rules.

**Rule of thumb.** Make something an entity type when users ask questions *about* it, or when it has its own attributes and relationships. If it's only a label on something else, make it an attribute, usually one with a [taxonomy](#taxonomy) or an enum.

### Attribute

**What it is.** A named, typed field on an entity type (or on a relationship). The types are **text**, **number**, **date**, **boolean**, **list** and **enum** (a fixed set of options). Each attribute can carry **extraction instructions** that tell the AI how to read it from documents, including the format you expect.

**What it's for.** Filters, sorts and facts in answers: dates, amounts, statuses, codes.

**Example.** `Contract.expiration_date` (date; "the date the contract term ends, not the signature date"), `Proposal.outcome` (enum: won / lost / pending).

### Relationship

**What it is.** An allowed, directed link between two entity types, named with an UPPER\_SNAKE verb. Relationships can have their own attributes.

**What it's for.** Following connections: who worked on what, for whom, under which contract. Multi-hop questions depend entirely on relationships being modeled and populated.

**Example.** `Employee -[WORKS_ON_PROJECT {role, start_date, hours}]-> Project`, `Contract -[IMPOSES]-> Obligation`.

### Taxonomy

**What it is.** A controlled, usually hierarchical list of terms of one type. Examples are Industry, Service Line and Skills. Each taxonomy type holds items, which can nest (`Healthcare > Providers`) and be active or inactive.

**What it's for.** Consistency. When an attribute's extraction instructions, or an enrichment rule, reference `@Industry`, the AI is given the taxonomy's active leaf values and must pick from them. You get "Healthcare > Payers" every time instead of "health insurance", "payer" and "HC payers". Chat also sees taxonomies when it plans queries, so users can filter by them.

**Where it comes from.** Often from another system: CRM picklists, an HR skills library, finance service-line codes, or industry standards such as NAICS. In that case it is **imported** as a CSV, not re-typed. Otherwise it's designed in the workshop.

**Example.** Northbridge's Industry taxonomy is imported from the Salesforce Industry picklist. ObligationType is designed with Legal.

**Where.** **Model & Define > Taxonomies** (CSV import on each taxonomy type's page). Built in [Workshop 4](/implementation/ws-taxonomies).

<Note>
  Taxonomies shape **extraction and enrichment**, not classification. They don't decide which artifact type a document is.
</Note>

### Artifact type

**What it is.** A family of documents that are read the same way, for example Statement of Work, Resume or Master Services Agreement. It is also called a **content type**, in the Admin Guide and older screens. It holds:

* **classification instructions**: how to recognise the family and tell it apart from look-alikes;
* the **subset of the ontology** to extract: entity types, attributes and relationships, with instruction overrides;
* per-entity **processing type** (Create Only / Create or Match / Match Only, plus Attach Document) and **required upstream** relationships (the parent is resolved first, so children match in its context);
* an **extraction policy**.

**What it's for.** It turns "read this PDF" into a precise, schema-bound task. A resume extracts the person and their skills. An SOW extracts the project, client, dates and value.

**Where.** **Model & Define > Artifact Types**. Includes **Test Ingestion** for trying samples. Built in [Workshop 5](/implementation/ws-artifact-types).

### Extraction policy

**What it is.** A setting on each artifact type for how deep, and how expensive, extraction is. It sets the mode (`full`, `metadata_and_snippet` or `metadata_only`), the model tier (large, medium or small), the validation pass (on or off) and an Excel override.

**What it's for.** Spending AI effort where it answers questions. Contracts get full extraction on a large model. Thousands of status reports might get metadata and a snippet.

**Where.** **Model & Define > Artifact Types > Ingestion extraction**. See [Extraction Policy](/admin-guide/extraction-policy).

### Data mapping

**What it is.** A column-by-column recipe that turns a structured source (a CSV, Excel or JSON file, or an API) into graph nodes, attributes and relationships. It uses **node mappings** (column → entity type and attribute, with create / match / create-or-match) and **relationship mappings** (which relationship, found by which columns at each end).

**What it's for.** Loading systems of record reliably, with no AI guesswork: the HR roster, CRM accounts, the finance project list, staffing assignments.

**Example.** `employees.csv` maps `work_email` to `Employee.email` (match property) and `manager_email` to `Employee -[REPORTS_TO]-> Employee`.

**Where.** **Model & Define > Data Mapping** (with **Auto-map with AI**). Built in [Workshop 6](/implementation/ws-data-mapping).

### Match key

**What it is.** The attribute used to decide that two records are the same thing: an ID, a code, an email or a canonical name.

**What it's for.** Joining sources. Structured loads match on it exactly. Documents match against the same nodes through their matching strategy. Choosing and cleaning match keys early prevents most duplicates.

**Example.** Employee → `email`. Project → `project_code`. Client → canonical Salesforce account `name`.

### Matching strategy

**What it is.** The rules for deciding whether an entity extracted from a document is the same as one already in the graph. It sets:

* the methods and their weights: **exact**, **fuzzy**, **phonetic**, **synonym** and **vector similarity**;
* **normalisation**: removing company suffixes, prefixes and punctuation;
* **filters**: time window, related parent;
* **thresholds**: auto-merge, AI decides, human review, new node.

There is a default strategy plus overrides per entity type.

**What it's for.** Preventing both duplicates ("Lakeshore Health" and "Lakeshore Health, Inc." as two clients) and false merges (two different people named "J. Smith" combined).

**Where.** **Process > Matching Strategies**. Built in [Workshop 7](/implementation/ws-identity-and-matching).

### Conflict resolution

**What it is.** The human review queues. **Classification reviews** hold files whose artifact type scored below 0.8. **Match reviews** hold entities whose match score fell in the human-review band, with Merge, Confirm New and Skip actions.

**What it's for.** Keeping a person in the loop exactly where the AI is unsure. The client super user usually owns these queues after go-live.

**Where.** **Process > Conflict Resolution**.

### Enrichment rule

**What it is.** An AI instruction that runs over nodes already in the graph and writes back an attribute, a new node, a relationship, or a node and a relationship. It can read the node, selected attributes, or its neighbours (`{related.Client.industry}`), and can be constrained to a taxonomy (`@Industry`). A relationship output links only to a target node that already exists; it doesn't create the target.

**What it's for.** Business rules that are true but not written in any one source. For example, tag each project's industry from its client, or infer the skills a project used from its scope.

**Where.** **Model & Define > Enrichment Rules**. Runs on demand or as a step in a [flow](#flow). Built in [Workshop 8](/implementation/ws-enrichment-rules).

### Client configuration

**What it is.** The firm's own identity: its canonical name, synonyms and branding.

**What it's for.** Extraction is told who "we" are, so "Northbridge", "NBC" and "Northbridge Consulting Group" aren't extracted as an outside company or a client.

**Where.** **Administer > Client Configuration**.

### Ontology compatibility

**What it is.** A check of every artifact type, data mapping, enrichment rule and matching strategy against the current ontology revision. Each one is **valid**, **stale** (a rename was applied; please review), **pending review** or **invalid** (it points at something deleted).

**What it's for.** Safe change. Only **invalid** blocks scans and rules. Check it after every ontology save.

**Where.** **Model & Define > Compatibility**. See [Ontology Compatibility](/admin-guide/ontology-compatibility).

## Pipeline concepts

### Connector

**What it is.** An authorised connection to a storage provider (Box, Google Drive or SharePoint) or an API.

**Where.** **Connect > Connectors**. Usually set up with the client's IT admin.

### Data source

**What it is.** What to read through a connector: the folders, recursion, file types, OCR, **folder filters** (which artifact types are allowed in which folders), and whether the source is documents or structured data.

**Where.** **Connect > Data Sources**.

### Flow

**What it is.** A scheduled or on-demand chain of **data jobs**: scan a data source, ingest, run enrichment rules, create structured jobs from folders.

**What it's for.** Keeping the graph current, for example a nightly scan followed by enrichment. It isn't the same as an [agent flow](#agent-flow).

**Where.** **Process > Flows** and **Flow Executions**.

### Lineage

**What it is.** The recorded origin of every node and relationship: which document, table row, enrichment rule or reviewer created or changed it, and when.

**What it's for.** Trust and debugging. Users trace a citation back to its source. Implementers find out why a wrong fact is in the graph.

**Where.** **View lineage** on citations and graph records. See [Graph Lineage](/admin-guide/graph-lineage).

### Graph evaluation

**What it is.** An AI judge that samples ingested documents per artifact type. It compares each document's text with what landed in the graph and reports missed or invented entities and relationships, wrong attributes, wrong relationship endpoints and schema violations.

**What it's for.** Measuring **extraction** quality per artifact type. It reports only and changes nothing. It doesn't test questions.

**Where.** **Process > Graph Evaluation**. See [Accuracy Validation](/implementation/accuracy-validation).

### Golden questions

**What it is.** The agreed list of real business questions, each with its expected answer and where the truth lives. It is gathered in [Workshop 1](/implementation/ws-questions-and-agents) and kept as `questions.md`, a Markdown file in the engagement's shared project folder. Start it from the [template](/implementation/templates#kickoff-and-discovery).

**What it's for.** It is the definition of success. It drives the ontology, and it is the test set for measuring **answer** accuracy before go-live and after every change. Scoring is a manual process today; see [Accuracy Validation](/implementation/accuracy-validation).

## Consumption concepts

### Assistant

**What it is.** A configured chat agent with a name, description, instructions, tools, model choices, and query instructions that encode business vocabulary.

**Where.** **AI & Agents > Agent Configuration**, with shared rules in **AI & Agents > AI Instructions**. Designed in [Workshop 10](/implementation/ws-assistants-and-flows).

### Agent flow

**What it is.** A multi-step AI workflow built on a visual canvas. Its blocks gather inputs, query the graph, run AI steps, pause for human review and produce files such as Word, PowerPoint, Excel or PDF. Users run it from chat.

**What it's for.** Repeatable jobs that produce an output, such as a past-performance write-up, a staffing shortlist or a contract renewal digest.

**Where.** **AI & Agents > Agent Flows**. Reviews wait in **Agent Inbox**. See [Agent Flows](/admin-guide/agent-flows).

### Document template

**What it is.** A branded Word, PowerPoint or Excel file, plus retrieval and output instructions, grouped by category (Case Study, Resume, Proposal and others).

**Where.** **Model & Define > Document Templates** and **Template Categories**.

### Persona

**What it is.** An optional audience profile, such as "Business development" or "New consultant", that tailors intake questions and answers. It is turned on per deployment.

**Where.** **AI & Agents > Personas**.

### Access control

**What it is.** Record-level visibility rules on the graph: which entity types are private or public, how access rolls up (for example, a contract under its client), and who is granted access (for example, members of the project team). It runs in disabled, shadow or enforce mode.

**What it's for.** Ethical walls and confidential data. Answers and citations respect it.

**Where.** **Administer > Access Control**. Designed in [Workshop 9](/implementation/ws-access-control).

### Knowledge-model copilot

<Info>
  In development. Availability depends on your release.
</Info>

**What it is.** An Admin Copilot mode for the Ontology and Artifact Types pages. You attach sample documents and tables along with `questions.md`, and it proposes ontology and artifact-type changes. It can also draft a starting ontology from samples. Each proposal is a card that a person reviews, applies and saves.

**What it's for.** Speeding up Workshops 3 and 5. The workshop reviews a draft instead of starting from a blank page. The manual path is always available.
