> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Workshop 4: Taxonomies

> Agree the controlled vocabularies the ontology uses, and import them from the systems that already own them.

This workshop decides the official lists of values Experio uses to tag things: industries, service lines, skills, obligation types. A [taxonomy](/implementation/key-concepts#taxonomy) gives the extraction model a closed list to choose from, so one document's "health care" and another's "Healthcare providers" land on the same value. Without it, a question like "healthcare clients" misses every client tagged with a variant spelling. Most of these lists already exist in a CRM, HR or finance system. The main job is to find the system of record, import from it, and agree who keeps it current.

## At a glance

| | |
| - | - |
| **Purpose** | Decide taxonomy vs. entity type for each controlled word, import existing lists, build missing ones, agree owners and refresh |
| **When** | Phase 2 Model, week 3, straight after [Workshop 3](/implementation/ws-ontology) |
| **Duration** | 90 minutes, plus import work by the super user afterwards |
| **Attendees** | Experio: FFE, Experio SMEs as needed. Client: PM, super user, the owner of each source list (sales ops, HR, finance), SMEs for any list built from scratch |
| **Inputs** | Saved ontology with attributes marked `@TaxonomyName`; list of controlled words from the question walk; exports of existing picklists |
| **Outputs** | One taxonomy type per list, imported or built; a synonyms list per taxonomy; an owner and refresh process per taxonomy |
| **Admin pages** | **Model & Define > Taxonomies** |

## Before the workshop

**FFE and Experio SMEs**

* From the Workshop 3 trace table, list every attribute that references a taxonomy and every controlled word the room raised.
* For each list, ask the client in advance: "Does this list exist in a system today? Who owns it?"
* Prepare a CSV template (see [CSV import format](#csv-import-format)) and send it with the homework.

**Client homework (list owners)**

* Export each existing list: CRM picklist values, the HR skills library, finance service-line codes, or the industry standard you use (NAICS, SIC).
* Include parent-child structure if the source has it, a description column if one exists, and a flag for retired values.

## Agenda

| Time | Segment | Output |
| - | - | - |
| 0:00 | Taxonomy vs. entity type: decide each candidate | Decision per candidate |
| 0:20 | Existing lists: system of record, export, gaps | Import plan per list |
| 0:40 | Build missing lists in the room | Draft hierarchy |
| 1:05 | Depth, synonyms, retired values | Agreed structure |
| 1:20 | Ownership and refresh | Owner and cadence per list |

## Running the workshop

### 1. Taxonomy or entity type?

For each controlled word, ask the room these questions in order.

| Question | If yes |
| - | - |
| Does it have its own attributes (a start date, an owner, a rate)? | Entity type |
| Does it have its own relationships (people have it, projects used it)? | Entity type |
| Do users ask questions about it as a thing ("which skills are we short on")? | Entity type |
| Is it only a label on something else, from an agreed list? | Taxonomy |
| Is the list three to five values that will never change and belong to one attribute? | Enum attribute in the ontology |

A word can be both. Northbridge models **Skill** as an entity type because people have skills with a proficiency and projects use them. The **Skills** taxonomy is the controlled list those Skill names come from, so extraction writes "Azure Data Platform" and not "Azure data stuff".

### 2. Find the system of record

Taxonomies frequently exist already. **Import them; do not re-type them.** Re-typed lists drift from the source within weeks and nobody owns them.

| Typical list | Usual system of record |
| - | - |
| Industry, account type, region | CRM picklists (Salesforce, Dynamics) |
| Skills, job levels, locations | HR system skills library (Workday, SuccessFactors) |
| Service lines, cost centres, practice codes | Finance or ERP codes |
| Industry classification | NAICS, SIC or another published standard |
| Document categories, matter types | Document management or legal system |

Ask: "If someone adds a new industry tomorrow, where do they add it?" That system is the source.

### 3. Build what does not exist

When no list exists, build it in the room with the SME. Keep it short, mutually exclusive, and written in the words the documents use. Ask: "If you had to put every item of this kind into one bucket, what are the buckets?" Then test it: read three real examples and have the room place each one.

### 4. Decide depth, synonyms and retired values

* **Depth.** Two or three levels is usually right. Extraction chooses only from **leaf** values, so the leaves must be specific enough to be useful and few enough to choose between reliably. Chat sees the whole hierarchy, so a question about a parent ("healthcare clients") can include all its leaves.
* **Synonyms.** Collect the variants people and documents use ("HC", "health care", "provider systems"). See the note below on where they are used today.
* **Retired values.** Keep them but mark them **inactive**. Inactive values drop out of extraction and chat prompts, and existing nodes keep their old value for history.

<Note>
  Each taxonomy item has a synonyms field, but in the current release the Taxonomies page does not show it, and synonyms are not included when a taxonomy is sent to the extraction model or chat. Keep the agreed synonyms list in the workshop notes, and put the variants that matter into the attribute's extraction instruction, for example "Treat 'HC' and 'health care' as healthcare values." Descriptions are also not sent to the model.
</Note>

### 5. Agree ownership and refresh

For each taxonomy, record: owner (a person, not a team), source system, refresh trigger (source list changes, or a fixed cadence such as quarterly), and who re-imports. Re-importing merges, so a refresh is a routine task for the super user.

## How taxonomies are used

* **Extraction.** In an attribute's extraction instruction (on the ontology or overridden on an artifact type), `@Industry` expands to the list of **active leaf values** of the Industry taxonomy.
* **Enrichment rules.** The same `@TaxonomyName` expansion works in enrichment prompts ([Workshop 8](/implementation/ws-enrichment-rules)).
* **Chat.** Query generation sees all active taxonomies with their hierarchy, so it can turn "healthcare" into the matching leaf values.
* **Not classification.** Taxonomies are not used when Experio decides a document's artifact type.

<Warning>
  `@` captures letters, digits, spaces, dashes and underscores until the next other character. Write `Choose one value from @Industry.` (the full stop ends the tag), or put the tag at the end of a line. `from @Industry and nothing else` looks up a taxonomy called "Industry and nothing else" and expands to an empty list. Name taxonomy types as one PascalCase word (`ObligationType`, not `Obligation Type`) so the tag matches.
</Warning>

## CSV import format

Upload a CSV on the taxonomy type's page. Import always **merges**: new items are created, missing parents are created along the path, and existing items are matched by name under the same parent. Nothing is deleted.

| Column | Required | Meaning |
| - | - | - |
| `path` | One of `path` or `level…` | Full path with `>` separators, for example `Healthcare > Providers` |
| `level1`, `level2`, … | One of `path` or `level…` | One column per level; blank trailing cells are fine |
| `description` | No | Stored on the last item in the row |
| `status` | No | `active` (default) or `inactive`, applied to the last item in the row |

Rules checked in the code:

* The first row is treated as a header only if it contains a `path`, `description` or `status` column. A file with only `level1,level2` headers would import the header row as items, so **always include a `description` or `status` column** with level columns.
* Item names cannot contain commas or `>`. Rows with a comma in a name are rejected. Rewrite "Oil, Gas & Mining" as "Oil Gas & Mining" or "Oil and Gas".
* On re-import, the `description` and `status` of each row's last item are set to the file's values. A blank description column clears existing descriptions, and a missing `status` column sets items back to active. Keep both columns in refresh files.

## Worked example: Northbridge Consulting

| Taxonomy | Source | Owner | Refresh |
| - | - | - | - |
| **Industry** | Salesforce Account Industry picklist | Sales ops | Re-import when the picklist changes |
| **ServiceLine** | Finance service-line codes | Finance/PMO | Quarterly, with the chart of accounts review |
| **Skills** | Workday skills library (\~180 skills) | HR | Quarterly |
| **ObligationType** | Built in the workshop with Helen Park | Helen Park | On request |

Dana confirmed the Industry picklist in Salesforce is two levels. Sales ops exported it and Marcus shaped it into a `path` file:

```csv theme={null}
path,description,status
Healthcare > Providers,"Hospitals, health systems, physician groups",active
Healthcare > Payers,Health insurers and managed care,active
Healthcare > Life Sciences,Pharma and medical devices,active
Financial Services > Banking,Retail and commercial banks,active
Financial Services > Insurance,Property and life insurers,active
Financial Services > Capital Markets,Asset managers and brokers,active
Public Sector > Federal,US federal agencies,active
Public Sector > State & Local,State and municipal government,active
```

Commas are allowed in the quoted `description`, not in item names. Finance supplied service lines with their codes, so Marcus used level columns and kept the code in the description:

```csv theme={null}
level1,level2,description,status
Strategy,Growth Strategy,Code STR-GS,active
Strategy,M&A,Code STR-MA,active
Technology,Cloud Migration,Code TEC-CM,active
Technology,Data & Analytics,Code TEC-DA,active
Technology,Cybersecurity,Code TEC-CY,active
Operations,Supply Chain,Code OPS-SC,active
Operations,Process Excellence,Code OPS-PE,active
```

No obligation list existed, so Helen built **ObligationType** in the room from ten MSAs: Reporting, Insurance, Confidentiality, Staffing, Service Levels, Data Protection. It is flat (one level), so every value is a leaf.

The ontology then references them in extraction instructions:

```text theme={null}
Client.industry: The client's industry. Choose one value from @Industry.
Project.service_line: The service line that delivered the work. Choose one value from @ServiceLine.
Obligation.obligation_type: The kind of obligation. Choose one value from @ObligationType.
```

When Salesforce later split Banking into Retail Banking and Commercial Banking, sales ops re-exported, Marcus added the two new rows and changed the Banking row to `inactive`, and re-imported. Nothing else changed.

## Entering it in Experio

<Steps>
  <Step title="Create the taxonomy type">
    Go to **Model & Define > Taxonomies** and add a new taxonomy type. Enter the display name exactly as the ontology will reference it (for example `Industry`) and a short description.
  </Step>

  <Step title="Import or build the items">
    Open the type and click **Upload CSV** to import, or add items and sub-items by hand for a list built in the workshop.
  </Step>

  <Step title="Check the tree">
    Expand the tree and confirm the hierarchy, names and leaf values. The import message reports created, updated and skipped rows; fix and re-upload any rejected rows.
  </Step>

  <Step title="Reference it from the ontology">
    In **Model & Define > Ontology**, add `@TaxonomyName` to the extraction instructions of the attributes that use it. Test with an artifact type's **Test Ingestion** in [Workshop 5](/implementation/ws-artifact-types).
  </Step>
</Steps>

Detail: [Taxonomies](/admin-guide/taxonomies). That page predates this guide and says taxonomies are used during classification; they are not.

## Common pitfalls

* **Re-typing a list that exists elsewhere.** It drifts from the source and has no owner. Export and import.
* **Parents as values.** Extraction picks leaves only. If "Healthcare" must be a value on its own, it needs no active children, or add a leaf such as "Healthcare > General".
* **Too deep or too long.** Hundreds of fine-grained leaves make the model guess. Trim to what questions distinguish.
* **Tag typos.** `@Industries` when the type is `Industry` expands to an empty list, silently.
* **Deleting retired values.** Deleting an item also deletes its children. Mark it inactive instead.
* **No refresh owner.** The first time the CRM list changes, the taxonomy is out of date.

## Exit criteria

* [ ] Every controlled word from Workshop 3 decided as taxonomy, entity type, or enum
* [ ] System of record and owner named for every imported taxonomy
* [ ] All taxonomies imported or built, with leaf values checked in the tree
* [ ] Taxonomy type names match the `@` references in the ontology exactly
* [ ] Synonyms list captured, and important variants added to extraction instructions
* [ ] Retired values marked inactive, not deleted
* [ ] Refresh process written down and the super user has done one re-import

## Next

[Workshop 5: Artifact Types](/implementation/ws-artifact-types) groups the documents and decides what to extract from each family.
