> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Workshop 3: Ontology

> Turn the golden questions into entity types, attributes, relationships and keys.

This workshop decides the shape of the knowledge graph: which things Experio tracks, what it records about each one, and how they connect. That shape is the [ontology](/implementation/key-concepts#ontology). Extraction can only fill in what the ontology defines, and chat can only answer questions whose nouns, verbs and filters exist in it. If a golden question needs "who led the project" and the ontology has no way to record a role on an assignment, no prompt tuning will fix the answer later. You build the ontology from the questions, not from the documents.

## At a glance

| | |
| - | - |
| **Purpose** | Agree the entity types, attributes, relationships, match keys and vector indexes that the golden questions need |
| **When** | Phase 2 Model, week 2–3. Run before Workshops 4–7 |
| **Duration** | 3 hours, plus a 60-minute follow-up after the first Test Ingestion runs |
| **Attendees** | Experio: FFE (facilitates), Experio SMEs as needed. Client: PM, super user, one SME per use case, the owner of each structured source |
| **Inputs** | Golden questions with expected answers ([Workshop 1](/implementation/ws-questions-and-agents)), data inventory with keys ([Workshop 2](/implementation/ws-data-inventory)), 2–3 sample documents per document family |
| **Outputs** | Ontology diagram, entity/attribute/relationship table, question-to-model trace, match key per entity type, list of taxonomies to build, a saved ontology revision |
| **Admin pages** | **Model & Define > Ontology**, **Model & Define > Compatibility** |

## Before the workshop

**FFE and Experio SMEs**

* Print or share the golden questions from Workshop 1, grouped by use case.
* Open **Model & Define > Ontology** and review the default ontology. Note which types fit the client (Client, Employee, Project, Proposal, Contract, Skill often do) and which do not (for example TimeEntry or ProductCatalog for a consulting firm).
* Pre-mark each question: underline nouns, circle verbs, box filters and sorts. You will redo this with the room, but a draft keeps the session moving.
* Pull the column headers of each structured export from the data inventory. The keys there (`employee_id`, `project_code`, `account_id`) become attributes and match keys.

**Client homework (super user)**

* Confirm the golden questions are final for the pilot. Adding questions later is fine; changing their meaning mid-workshop is not.
* Bring the headers of every structured export and one real example row.
* Bring any existing picklists you already know about (industries, service lines, skills). They feed [Workshop 4](/implementation/ws-taxonomies).

## Agenda

| Time | Segment | Output |
| - | - | - |
| 0:00 | Why the ontology matters; walk one question end to end | Shared vocabulary |
| 0:15 | Question walk: nouns, verbs, filters, controlled words | Raw candidate list |
| 1:00 | Consolidate entity types; compare with the default ontology | 8–15 entity types |
| 1:30 | Break | |
| 1:40 | Attributes: types, extraction instructions, formats | Attribute table |
| 2:10 | Relationships: direction, names, relationship attributes | Relationship table |
| 2:35 | Keys and vector indexes | Match key per type, index list |
| 2:50 | Trace check: every question maps to the model; scope cut | Trace table, parking lot |

## Running the workshop

### 1. Walk each question

Put one question on screen at a time and mark it up with the room.

| Mark | Becomes | Question to ask the room |
| - | - | - |
| Nouns ("projects", "clients", "consultants") | [Entity types](/implementation/key-concepts#entity-type) | "Is this a thing we want to list, count or look up by name?" |
| Verbs ("delivered for", "led", "worked on") | [Relationships](/implementation/key-concepts#relationship) | "Which two things does this connect, and which way does it read?" |
| Filters and sorts ("since 2023", "below \$1M", "in the next 90 days") | [Attributes](/implementation/key-concepts#attribute) with a type | "Would you filter or sort by this? Is it a date, a number, yes/no, or a choice from a list?" |
| Controlled words ("healthcare", "cloud migration", "reporting obligation") | [Taxonomies](/implementation/key-concepts#taxonomy) or enum attributes | "Is there an official list of these values somewhere?" |

Write every candidate down, even duplicates. Consolidate afterwards.

### 2. Consolidate entity types

* Use **singular PascalCase** names: `Client`, `Project`, `Employee`, `Obligation`. Not `Clients`, not `client_account`.
* Merge synonyms into one type. "Customer", "account" and "client" are one `Client`. Record the other words as synonyms on the entity type; chat uses them to map informal terms to the right type.
* Start from the default ontology where it fits. Keep its types and attributes when they match, rename when only the name differs (for example `Division` to `Practice`), and delete what no question needs.
* Aim for **8–15 entity types** for a pilot. More than that usually means you are modelling the documents rather than the questions.

Ask: "Which question breaks if we drop this type?" If none does, drop it or put it on the parking list.

### 3. Decide attributes

For each entity type, list only the attributes a question filters on, sorts on, returns, or needs as a key.

| Type | Use for | Example |
| - | - | - |
| Text | Names, descriptions, codes | `project_code`, `description` |
| Number | Anything you compare or add up | `contract_value`, `liability_cap` |
| Date | Anything with "since", "before", "expiring" | `start_date`, `expiration_date` |
| Boolean | Yes/no flags | `is_active` |
| List | Several values on one node | `aliases` |
| Enum (with options) | A short fixed list that belongs to this attribute only | `outcome`: won / lost / pending |

Write an **extraction instruction** for every attribute a document will fill. Say where the value usually appears, what to do when it is missing, and the format you want. Put the format in the instruction itself, for example "Format: YYYY-MM-DD" or "Number in US dollars, no currency symbol". Long lists of allowed values belong in a taxonomy; reference it with `@TaxonomyName` (see [Workshop 4](/implementation/ws-taxonomies)).

<Note>
  The visual editor's type list offers Text, Number, Date and List. Set **boolean** and **enum** attributes in the **JSON** view; an enum needs an `options` array. Enum options are validated when you save but are not sent to the extraction model, so also list the allowed values in the extraction instruction ("One of: MSA, SOW, amendment."). Values are stored as text.
</Note>

### 4. Decide relationships

* Name relationships as **UPPER\_SNAKE verbs** that read left to right: `Employee WORKS_ON_PROJECT Project`, `Contract IMPOSES Obligation`.
* Model **one direction only**. Chat can follow a relationship either way, so do not add inverse pairs like `HAS_EMPLOYEE` and `WORKS_FOR`.
* Put facts about the connection on the relationship, not on either end. The role someone played on a project belongs on `WORKS_ON_PROJECT {role, start_date, end_date, hours}`, because the same person has different roles on different projects.
* Mark a relationship **Required upstream relationship** only when the child cannot be created correctly without its parent. The flag treats the relationship's **source** as the parent, so the relationship must point parent to child, and the source type must have a vector index. See [Workshop 5](/implementation/ws-artifact-types) for how artifact types use it.

Ask: "When this appears in a document, which end is the thing the document is about?"

### 5. Decide keys and vector indexes

For each entity type, agree the **[match key](/implementation/key-concepts#match-key)**: the attribute that identifies one real-world thing across sources. Record it now; you enter it later in data mappings ([Workshop 6](/implementation/ws-data-mapping)) and matching strategies ([Workshop 7](/implementation/ws-identity-and-matching)). The ontology itself has no "key" field, but the key must exist as an attribute.

Turn on a **vector index** for entity types that users mention by name in questions, and for types that documents must match by similarity. Chat resolves names in a question ("Lakeshore", "Bob Chen") through exact matches and these indexes, and document matching can only use Vector Similarity on an indexed type. Small closed lists (Practice) rarely need one.

### 6. Trace and cut

Fill in the question-to-model trace table (see the worked example). Every golden question must map to types, relationships and attributes that exist. Anything in the model that no question uses goes to a parking list for after the pilot.

## Worked example: Northbridge Consulting

Marcus Lee (super user), Dana Ortiz, Sam Whitfield and Helen Park walked the nine golden questions. They kept Client, Employee, Project, Proposal, Contract, Skill and Obligation from the default ontology, renamed Division to Practice, and deleted the rest.

```mermaid theme={null}
flowchart LR
  Proposal -- SUBMITTED_TO --> Client
  Proposal -- RESULTED_IN --> Project
  Project -- FOR_CLIENT --> Client
  Project -- DELIVERED_BY --> Practice
  Project -- USED_SKILL --> Skill
  Employee -- "WORKS_ON_PROJECT {role, start_date, end_date, hours}" --> Project
  Employee -- "HAS_SKILL {proficiency}" --> Skill
  Employee -- MEMBER_OF --> Practice
  Employee -- REPORTS_TO --> Employee
  Contract -- WITH_CLIENT --> Client
  Contract -- GOVERNS --> Project
  Contract -- AMENDS --> Contract
  Contract -- IMPOSES --> Obligation
```

| Entity type | Attributes (type) | Match key | Vector index |
| - | - | - | - |
| Client | name (text), account\_id (text), industry (text, `@Industry`), region (text) | `name` (canonical Salesforce name) | Yes, name |
| Project | name, project\_code (text), start\_date, end\_date (date), contract\_value (number), service\_line (text, `@ServiceLine`), description (text) | `project_code` | Yes, name |
| Employee | name, employee\_id, email, title, level, location (text) | `email` (structured); name for documents | Yes, name |
| Practice | name (text) | `name` | No |
| Skill | name (text, from `@Skills`), category (text) | `name` | Yes, name |
| Proposal | name (text), submitted\_date (date), outcome (enum: won / lost / pending), value (number) | `name` | No |
| Contract | name (text), contract\_type (enum: MSA / SOW / amendment), effective\_date, expiration\_date (date), liability\_cap (number) | `name` | Yes, name |
| Obligation | name (short label), description (text), obligation\_type (text, `@ObligationType`), due\_date (date), frequency (enum) | none; always created in the context of its Contract | Optional, on description |

| Relationship | Attributes | Notes |
| - | - | - |
| Project FOR\_CLIENT Client | | Created from `projects.csv` |
| Employee WORKS\_ON\_PROJECT Project | role, start\_date, end\_date, hours | From `assignments.csv`; answers "who led" |
| Employee HAS\_SKILL Skill | proficiency | From resumes |
| Employee MEMBER\_OF Practice; Project DELIVERED\_BY Practice | | From HR and project exports |
| Employee REPORTS\_TO Employee | | From `employees.csv` manager email |
| Proposal SUBMITTED\_TO Client; Proposal RESULTED\_IN Project | | From proposals and CRM |
| Contract WITH\_CLIENT Client; Contract GOVERNS Project; Contract AMENDS Contract | | From SOWs and MSAs |
| Contract IMPOSES Obligation | | Required upstream: obligations resolve under their contract |
| Project USED\_SKILL Skill | | Created by an enrichment rule ([Workshop 8](/implementation/ws-enrichment-rules)) |

Sample extraction instructions Helen agreed for Contract:

```text theme={null}
expiration_date: The date the agreement ends. Look for "Term", "Expiration Date" or
"shall remain in effect until". If the term is stated as a duration, add it to the
effective date. Leave blank if the contract renews automatically with no end date.
Format: YYYY-MM-DD.

liability_cap: The maximum aggregate liability in the "Limitation of Liability" clause.
If the cap is a multiple of fees, leave blank and quote the clause in description.
Format: number in US dollars, no currency symbol or commas.
```

### Question to model trace

| Question | Entity types | Relationships | Attributes and taxonomies |
| - | - | - | - |
| Q1 Cloud migration projects for healthcare clients since 2023, and who led each | Project, Client, Employee | FOR\_CLIENT, WORKS\_ON\_PROJECT | Project.service\_line (`@ServiceLine`: Cloud Migration), Client.industry (`@Industry`: Healthcare leaves), Project.start\_date, WORKS\_ON\_PROJECT.role = "Engagement Lead" |
| Q4 Azure data platform experience and healthcare work | Employee, Skill, Project, Client | HAS\_SKILL, WORKS\_ON\_PROJECT, FOR\_CLIENT | Skill.name (`@Skills`: Azure Data Platform), Client.industry |
| Q7 SOWs expiring in the next 90 days, and for which clients | Contract, Client | WITH\_CLIENT | Contract.contract\_type = SOW, Contract.expiration\_date |
| Q8 Reporting obligations under the Lakeshore Health MSA | Contract, Client, Obligation | WITH\_CLIENT, IMPOSES | Contract.contract\_type = MSA, Client.name, Obligation.obligation\_type (`@ObligationType`: Reporting) |

Q7 showed that `expiration_date` must be a **date**, not text, or "next 90 days" cannot be computed. Q1 showed that "who led" needs `role` on the relationship, not a `lead` attribute on Project.

## Accelerate with AI

<Info>
  In development — availability depends on your release.
</Info>

The [knowledge-model copilot](/implementation/key-concepts#knowledge-model-copilot) can draft a starting ontology. Attach 3–5 sample documents, the structured export headers and the golden questions file (`questions.md`). It proposes entity types, attributes and relationships as cards. Treat the draft as pre-reading: review each card in the workshop, apply the ones the room agrees, and save on the Ontology page. The manual path above always works and is the one to use when the copilot is not in your release.

**Admin Copilot** (⌘J) is available today and can answer "how do I" questions about the ontology editor from these docs.

## Entering it in Experio

<Steps>
  <Step title="Open the editor">
    Go to **Model & Define > Ontology**. Use the **Canvas** view to add entity types and draw relationships, or the **Json** view to edit the schema directly (needed for boolean and enum attributes).
  </Step>

  <Step title="Add entity types">
    For each type, fill **Node Name**, **Semantic Intent** (one sentence on what it means) and **Synonyms** on the Details tab, then add attributes with type and extraction instructions on the **Attributes** tab.
  </Step>

  <Step title="Add relationships">
    Choose source entity, relation type and target entity. Add relationship attributes (for example `role`). Tick **Required upstream relationship** only where agreed.
  </Step>

  <Step title="Turn on vector indexes">
    Select the entity type and open the **Indexes** tab. Enable the name index (the expression defaults to `name`) and any attribute indexes on text or list attributes.
  </Step>

  <Step title="Save a revision">
    Click **Save**. If the change renames or deletes anything, a confirmation lists the impact. Each save creates a new revision you can compare or roll back from Revision History.
  </Step>

  <Step title="Check compatibility">
    Open **Model & Define > Compatibility** and fix anything **invalid** before the next scan.
  </Step>
</Steps>

### Revisions and compatibility

Every save publishes a revision, and Experio re-checks everything that depends on the ontology: artifact types, data mappings, enrichment rules and matching strategies. See [ontology compatibility](/implementation/key-concepts#ontology-compatibility).

| Change | Effect on dependents | Blocks ingestion? |
| - | - | - |
| Add a type, attribute or relationship | None | No |
| Rename | Names update automatically; dependents become **stale** for review | No (warning only) |
| Delete | Dependents that still reference it become **invalid** | Yes, until fixed and re-validated |

Deleting from the ontology does not delete nodes already in the graph. Early in the project, before artifact types and mappings exist, changes are cheap. After pilot ingestion, prefer renames and additions, and plan deletes together with the fixes to dependents. Detail: [Ontology](/admin-guide/ontology) and [Ontology Compatibility](/admin-guide/ontology-compatibility).

## Common pitfalls

* **Modelling the documents.** A `Resume` entity type with an `experience` blob answers nothing. Model the person, skills and projects the resume describes.
* **Facts on the wrong end.** A role, date range or proficiency that depends on the pair belongs on the relationship.
* **Dates and money as text.** Filters like "since 2023" and "below \$1M" need date and number types.
* **Enum values only in `options`.** The extraction model does not see them; repeat them in the instruction.
* **No key.** If two sources describe the same client and there is no agreed key, you will get duplicates. Decide keys now.
* **Inverse relationships.** `HAS_PROJECT` and `FOR_CLIENT` side by side confuse extraction and chat. Keep one.
* **Keeping unused default types.** Every extra type is something extraction may try to fill. Delete what no question needs, and remove any seeded artifact types that reference deleted types so they do not show as invalid.

## Exit criteria

* [ ] Every golden question appears in the trace table with types, relationships and attributes that exist in the saved ontology
* [ ] 8–15 entity types, singular PascalCase; relationships UPPER\_SNAKE, one direction each
* [ ] Every filter or sort attribute has the right type (date, number, boolean, enum)
* [ ] Extraction instructions written for every attribute documents will fill, including format and allowed values
* [ ] Match key agreed for every entity type and present as an attribute
* [ ] Vector indexes enabled on the types users name in questions
* [ ] Taxonomies needed (`@Industry`, `@ServiceLine` …) listed for Workshop 4
* [ ] Ontology saved; **Model & Define > Compatibility** shows no invalid configs
* [ ] Super user can explain the model to a colleague using the diagram

## Next

[Workshop 4: Taxonomies](/implementation/ws-taxonomies) builds or imports the controlled lists the ontology references.
