Skip to main content
This workshop decides the official lists of values Experio uses to tag things: industries, service lines, skills, obligation types. A taxonomy gives the extraction model a closed list to choose from, so one document’s “health care” and another’s “Healthcare providers” land on the same value. Without it, a question like “healthcare clients” misses every client tagged with a variant spelling. Most of these lists already exist in a CRM, HR or finance system. The main job is to find the system of record, import from it, and agree who keeps it current.

At a glance

Before the workshop

FFE and Experio SMEs
  • From the Workshop 3 trace table, list every attribute that references a taxonomy and every controlled word the room raised.
  • For each list, ask the client in advance: “Does this list exist in a system today? Who owns it?”
  • Prepare a CSV template (see CSV import format) and send it with the homework.
Client homework (list owners)
  • Export each existing list: CRM picklist values, the HR skills library, finance service-line codes, or the industry standard you use (NAICS, SIC).
  • Include parent-child structure if the source has it, a description column if one exists, and a flag for retired values.

Agenda

Running the workshop

1. Taxonomy or entity type?

For each controlled word, ask the room these questions in order. A word can be both. Northbridge models Skill as an entity type because people have skills with a proficiency and projects use them. The Skills taxonomy is the controlled list those Skill names come from, so extraction writes “Azure Data Platform” and not “Azure data stuff”.

2. Find the system of record

Taxonomies frequently exist already. Import them; do not re-type them. Re-typed lists drift from the source within weeks and nobody owns them. Ask: “If someone adds a new industry tomorrow, where do they add it?” That system is the source.

3. Build what does not exist

When no list exists, build it in the room with the SME. Keep it short, mutually exclusive, and written in the words the documents use. Ask: “If you had to put every item of this kind into one bucket, what are the buckets?” Then test it: read three real examples and have the room place each one.

4. Decide depth, synonyms and retired values

  • Depth. Two or three levels is usually right. Extraction chooses only from leaf values, so the leaves must be specific enough to be useful and few enough to choose between reliably. Chat sees the whole hierarchy, so a question about a parent (“healthcare clients”) can include all its leaves.
  • Synonyms. Collect the variants people and documents use (“HC”, “health care”, “provider systems”). See the note below on where they are used today.
  • Retired values. Keep them but mark them inactive. Inactive values drop out of extraction and chat prompts, and existing nodes keep their old value for history.
Each taxonomy item has a synonyms field, but in the current release the Taxonomies page does not show it, and synonyms are not included when a taxonomy is sent to the extraction model or chat. Keep the agreed synonyms list in the workshop notes, and put the variants that matter into the attribute’s extraction instruction, for example “Treat ‘HC’ and ‘health care’ as healthcare values.” Descriptions are also not sent to the model.

5. Agree ownership and refresh

For each taxonomy, record: owner (a person, not a team), source system, refresh trigger (source list changes, or a fixed cadence such as quarterly), and who re-imports. Re-importing merges, so a refresh is a routine task for the super user.

How taxonomies are used

  • Extraction. In an attribute’s extraction instruction (on the ontology or overridden on an artifact type), @Industry expands to the list of active leaf values of the Industry taxonomy.
  • Enrichment rules. The same @TaxonomyName expansion works in enrichment prompts (Workshop 8).
  • Chat. Query generation sees all active taxonomies with their hierarchy, so it can turn “healthcare” into the matching leaf values.
  • Not classification. Taxonomies are not used when Experio decides a document’s artifact type.
@ captures letters, digits, spaces, dashes and underscores until the next other character. Write Choose one value from @Industry. (the full stop ends the tag), or put the tag at the end of a line. from @Industry and nothing else looks up a taxonomy called “Industry and nothing else” and expands to an empty list. Name taxonomy types as one PascalCase word (ObligationType, not Obligation Type) so the tag matches.

CSV import format

Upload a CSV on the taxonomy type’s page. Import always merges: new items are created, missing parents are created along the path, and existing items are matched by name under the same parent. Nothing is deleted. Rules checked in the code:
  • The first row is treated as a header only if it contains a path, description or status column. A file with only level1,level2 headers would import the header row as items, so always include a description or status column with level columns.
  • Item names cannot contain commas or >. Rows with a comma in a name are rejected. Rewrite “Oil, Gas & Mining” as “Oil Gas & Mining” or “Oil and Gas”.
  • On re-import, the description and status of each row’s last item are set to the file’s values. A blank description column clears existing descriptions, and a missing status column sets items back to active. Keep both columns in refresh files.

Worked example: Northbridge Consulting

Dana confirmed the Industry picklist in Salesforce is two levels. Sales ops exported it and Marcus shaped it into a path file:
Commas are allowed in the quoted description, not in item names. Finance supplied service lines with their codes, so Marcus used level columns and kept the code in the description:
No obligation list existed, so Helen built ObligationType in the room from ten MSAs: Reporting, Insurance, Confidentiality, Staffing, Service Levels, Data Protection. It is flat (one level), so every value is a leaf. The ontology then references them in extraction instructions:
When Salesforce later split Banking into Retail Banking and Commercial Banking, sales ops re-exported, Marcus added the two new rows and changed the Banking row to inactive, and re-imported. Nothing else changed.

Entering it in Experio

1

Create the taxonomy type

Go to Model & Define > Taxonomies and add a new taxonomy type. Enter the display name exactly as the ontology will reference it (for example Industry) and a short description.
2

Import or build the items

Open the type and click Upload CSV to import, or add items and sub-items by hand for a list built in the workshop.
3

Check the tree

Expand the tree and confirm the hierarchy, names and leaf values. The import message reports created, updated and skipped rows; fix and re-upload any rejected rows.
4

Reference it from the ontology

In Model & Define > Ontology, add @TaxonomyName to the extraction instructions of the attributes that use it. Test with an artifact type’s Test Ingestion in Workshop 5.
Detail: Taxonomies. That page predates this guide and says taxonomies are used during classification; they are not.

Common pitfalls

  • Re-typing a list that exists elsewhere. It drifts from the source and has no owner. Export and import.
  • Parents as values. Extraction picks leaves only. If “Healthcare” must be a value on its own, it needs no active children, or add a leaf such as “Healthcare > General”.
  • Too deep or too long. Hundreds of fine-grained leaves make the model guess. Trim to what questions distinguish.
  • Tag typos. @Industries when the type is Industry expands to an empty list, silently.
  • Deleting retired values. Deleting an item also deletes its children. Mark it inactive instead.
  • No refresh owner. The first time the CRM list changes, the taxonomy is out of date.

Exit criteria

  • Every controlled word from Workshop 3 decided as taxonomy, entity type, or enum
  • System of record and owner named for every imported taxonomy
  • All taxonomies imported or built, with leaf values checked in the tree
  • Taxonomy type names match the @ references in the ontology exactly
  • Synonyms list captured, and important variants added to extraction instructions
  • Retired values marked inactive, not deleted
  • Refresh process written down and the super user has done one re-import

Next

Workshop 5: Artifact Types groups the documents and decides what to extract from each family.