> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Workshop 5: Artifact Types

> Group the documents into families and decide how Experio recognises each one and what it extracts.

This workshop decides how Experio treats each kind of document. An [artifact type](/implementation/key-concepts#artifact-type) (the admin guide calls it a content type) tells the classifier how to recognise a document family, and tells extraction which part of the ontology to fill from it. Two decisions carry most of the accuracy: how you tell look-alike documents apart (a SOW from an MSA from an amendment), and which entities a document may create versus only match. You make both decisions with real documents on screen, using Test Ingestion as you go.

## At a glance

| | |
| - | - |
| **Purpose** | Define each document family: classification instructions, extraction subset, processing type per entity, extraction policy, folder filters |
| **When** | Phase 2 Model, week 3–4, after [Workshop 3](/implementation/ws-ontology) and [Workshop 4](/implementation/ws-taxonomies) |
| **Duration** | 3 hours for 5–6 families; add 30 minutes per extra family |
| **Attendees** | Experio: FFE (drives the admin screens), Experio SMEs as needed. Client: PM, super user, the SME who owns each family (BD, Legal, HR, PMO), IT admin for folder structure |
| **Inputs** | Saved ontology and taxonomies; data inventory with folder structure; **3–5 real sample documents per family**, including awkward ones |
| **Outputs** | One artifact type per family, tested; a spec sheet per family; folder filters mapped to allowed types; extraction policy per type |
| **Admin pages** | **Model & Define > Artifact Types**, **Connect > Data Sources** (folder filters), **Process > Conflict Resolution** (classification reviews) |

## Before the workshop

**FFE and Experio SMEs**

* From the data inventory, list the document families and where each lives. Propose a first grouping.
* Create a draft artifact type per family with a name and one-line classification instruction, so the workshop starts with something to test.
* Upload the sample documents to a File Upload data source and run a job, so they are parsed and available to the AI Assistant.

**Client homework (SMEs)**

* Collect 3–5 real documents per family: a typical one, an old-format one, a short one, and one that is easy to confuse with another family (for example an SOW that contains full contract terms).
* Note how you tell the family apart from its look-alikes. "Our SOWs always reference an MSA and have a project code in the header" is exactly what the classifier needs.

## Agenda

| Time | Segment | Output |
| - | - | - |
| 0:00 | How classification and extraction work; the 0.8 threshold | Shared understanding |
| 0:15 | Group the corpus into families; decide what is out of scope | Family list, "None" list |
| 0:35 | Per family: classification instructions; Test Ingestion **Classify Only** | Instructions that separate look-alikes |
| 1:30 | Break | |
| 1:40 | Per family: entities, attributes, relationships, processing types | Extraction subset |
| 2:20 | Test Ingestion **Classify & Extract** on samples | Corrections |
| 2:40 | Extraction policy per family; folder filters | Cost/accuracy decisions, filter map |

## Running the workshop

### 1. Group the corpus into families

A family is a set of documents that serve the same purpose and answer the same questions. Group by purpose, not by file format or folder. Ask: "Which golden questions does this kind of document answer?" A family that answers no question does not need an artifact type; it will be classified as None and skipped.

### 2. Write classification instructions

The classifier scores every allowed artifact type from 0 to 1 against its classification instructions. A document is assigned a type only if its best score reaches **0.8**. This threshold is fixed. Below it, a classification review goes to **Process > Conflict Resolution** and the file waits for a person to choose. A result of None, or a type not allowed for that folder, means the file is skipped.

Good instructions have three parts:

1. **What it is**: purpose and typical titles.
2. **How to recognise it**: headings, clauses, layout, who signs it.
3. **How to tell it from look-alikes**: name the confusable families and the deciding feature.

Run **Classify Only** on each sample as you write. If a sample scores 0.6 for two types, the instructions do not separate them yet. Ask the SME: "What is the first thing you look at to tell these apart?"

### 3. Choose what to extract

Pick the subset of the ontology this family can actually supply. Ask: "If this document were the only source, which golden-question facts would it give us?" Extract only those.

* Select the **entities** and, per entity, the **attributes**. Extraction instructions default from the ontology; override them here when this family words things differently.
* Select the **relationships** between the chosen entities, with their attributes (for example `role`).
* Mark attributes **required** only if the document always has them.

### 4. Decide the processing type per entity

| Processing type | What it does | Use when |
| - | - | - |
| **Create Only** | Always creates a new node | Things that exist only in this document (an obligation clause) |
| **Create or Match** | Runs [matching](/implementation/key-concepts#matching-strategy); merges with an existing node or creates a new one | Things that may already exist from other documents or structured loads (clients, people) |
| **Match Only** | Updates an existing node; never creates one | Things a system of record owns (projects from finance) |
| **Attach Document** | Links the document to this entity; one entity per artifact type | The entity the document is "about" (the contract, the person) |

Load structured master data first ([Workshop 6](/implementation/ws-data-mapping)), then let documents Create or Match or Match Only against it. Matching detail is [Workshop 7](/implementation/ws-identity-and-matching).

Mark a relationship **Required upstream dependency** when the child should resolve in the context of its parent. The relationship's source is the parent, and the parent type needs a vector index (set in [Workshop 3](/implementation/ws-ontology)).

### 5. Set the extraction policy

The [extraction policy](/implementation/key-concepts#extraction-policy) is a cost and accuracy trade-off per family.

| Setting | Options | Guidance |
| - | - | - |
| Mode | `full` / `metadata_and_snippet` / `metadata_only` | `full` for families that answer questions; snippet or metadata for high-volume, low-value files |
| Model tier | large / medium / small | Large for contracts and anything legal; medium often suffices for resumes and case studies |
| Validation pass | on / off | On where missing an entity is costly (obligations); off for simple, short documents |
| Excel override | same as default, or a different mode | Usually `metadata_only` for exports and pricing workbooks |

In the two metadata modes Experio skips entity extraction. It writes one shell node of the attached (or first) entity type, named after the file, plus the snippet when there is one, and no relationships. The parsed text is still stored on the Document node for chat. A cost guard also drops very large files to metadata-only. See [Extraction Policy](/admin-guide/extraction-policy).

### 6. Map folders to allowed types

Each data source filter can list the artifact types allowed for the files it matches. The allowed set is the union across matching filters, or every type if no filter lists any. Narrowing the set is the cheapest accuracy gain there is: a file in the Resumes library cannot be misclassified as an MSA. Ask the IT admin: "Is each folder dedicated to one kind of document, or mixed?"

## Worked example: Northbridge Consulting

| Artifact type | Entities (processing type) | Key relationships | Policy | Folder filter |
| - | - | - | - | - |
| Proposal | Proposal (Create or Match, Attach), Client (Create or Match) | SUBMITTED\_TO | full, large, validation off | Engagements |
| Statement of Work | Contract (Create or Match, Attach), Client (Create or Match), Project (Match Only) | WITH\_CLIENT, GOVERNS, AMENDS | full, large, validation on | Engagements |
| Master Services Agreement | Contract (Create or Match, Attach), Client (Create or Match), Obligation (Create Only) | WITH\_CLIENT, IMPOSES (required upstream) | full, large, validation on | Engagements |
| Resume | Employee (Create or Match, Attach), Skill (Create or Match) | HAS\_SKILL {proficiency} | full, medium, validation off | Resumes |
| Case Study | Project (Match Only, Attach), Client (Create or Match) | FOR\_CLIENT | full, medium, validation off | Engagements |
| Status Report | Project (Match Only, Attach) | none | metadata\_and\_snippet, small; Excel metadata\_only | Engagements |

Everything else (templates, invoices, internal memos) is classified as None and skipped. The Resumes library filter allows **Resume** only; the Engagements filter allows Proposal, Statement of Work, Master Services Agreement, Case Study and Status Report.

### Spec: Statement of Work

Helen Park brought four SOWs, one amendment and one MSA. The first draft scored the amendment 0.85 as a SOW and the MSA 0.7 as a SOW, so the room added look-alike rules.

```text theme={null}
A Statement of Work (SOW) defines the scope, deliverables, timeline, staffing and fees
for one engagement under an existing agreement. Titles include "Statement of Work",
"SOW", "Work Order" and "Change Order". Northbridge SOWs usually show a project code
such as NB-2024-117 in the header and state that they are issued under a Master
Services Agreement.

Classify as SOW:
- Amendments and change orders that modify a SOW's scope, dates or fees.

Do NOT classify as SOW:
- Master Services Agreements: general terms (liability, indemnity, confidentiality)
  with no single project scope, even if they include a sample SOW as an exhibit.
- Proposals: written before award, offering work rather than agreeing it; no
  signature blocks for both parties.
```

Extraction for the SOW family:

| Entity / relationship | Attribute | Extraction instruction (override where shown) |
| - | - | - |
| Contract | name | "Title and number of the SOW, for example 'SOW 3 – Cloud Migration'." |
| Contract | contract\_type | Override: "SOW for a statement of work; amendment for an amendment or change order. One of: SOW, amendment." |
| Contract | effective\_date, expiration\_date | From ontology (Format: YYYY-MM-DD) |
| Client | name | "The client party, not Northbridge. Use the legal name in the preamble." |
| Project | project\_code | "The Northbridge project code in the header, format NB-YYYY-NNN." |
| Contract AMENDS Contract | | "Only for amendments: the SOW being amended." |

Project is **Match Only**: a SOW must never create a project that finance does not know. Northbridge is kept out of Client automatically, because the firm name and its synonyms from Client Configuration are sent with every extraction.

### Spec: Resume

Sam Whitfield's samples came in three layouts. The classification instruction was short, because the Resumes folder allows nothing else:

```text theme={null}
A resume or CV describing one Northbridge employee's experience, skills, education
and certifications. Usually has the person's name at the top and sections such as
"Experience", "Skills" and "Education". Not a project team list or staffing plan,
which names several people.
```

| Entity / relationship | Attribute | Extraction instruction |
| - | - | - |
| Employee | name, title | "The person the resume is about. Use the name at the top." |
| Employee | email | "Work email if shown; leave blank otherwise." |
| Skill | name | "Each technical or domain skill. Choose one value from @Skills." |
| Employee HAS\_SKILL Skill | proficiency | "Expert, Advanced or Working, from the resume's own wording. Leave blank if not stated." |

The room decided **not** to extract projects from resumes. Resume project names ("Lakeshore cloud work") cannot be matched to project codes, and `assignments.csv` already holds the truth for who worked on what.

## Accelerate with AI

**Available today: Artifact Types > AI Assistant.** Open an artifact type, open the AI Assistant, and choose up to 10 sample files that a job has already parsed. It drafts the extraction schema against the existing ontology, and you refine it by chatting ("drop Obligation, add liability\_cap"). It needs an ontology first. Use it to produce the first draft before the workshop, then correct it with the SMEs using Test Ingestion.

<Info>
  In development — availability depends on your release.
</Info>

The [knowledge-model copilot](/implementation/key-concepts#knowledge-model-copilot) takes sample documents plus the golden questions (`questions.md`) and proposes artifact-type changes as cards, which a person reviews, applies and saves on the Artifact Types page. The manual path always works.

## Entering it in Experio

<Steps>
  <Step title="Create the artifact type">
    Go to **Model & Define > Artifact Types** and create a type. Enter the name and classification instructions, and set **Ingestion extraction** (mode, model tier, validation pass, Excel mode).
  </Step>

  <Step title="Add entities and relationships">
    Add each entity with its processing type and, for one entity, **Attach document to this entity**. Choose attributes and override extraction instructions where needed. Add relationships, their attributes and any required upstream dependency.
  </Step>

  <Step title="Test">
    Click **Test Ingestion**, upload a sample, and run **Classify Only** first, then **Classify & Extract**. Review classification scores, entities, relationships and raw JSON. Use **Max Pages** to test long documents quickly.
  </Step>

  <Step title="Map folder filters">
    In **Connect > Data Sources**, edit each filter and choose the artifact types allowed for the files it matches.
  </Step>
</Steps>

Detail: [Content Types](/admin-guide/content-types), [Extraction Policy](/admin-guide/extraction-policy), [Data Sources](/admin-guide/data-sources).

## Common pitfalls

* **Instructions that only describe.** Without "how to tell it from X", look-alikes score close together and land in review.
* **Testing only typical samples.** Bring the awkward ones; they are what fails in production.
* **Extracting everything.** Each extra entity is cost and a chance of noise. Extract what the questions need.
* **Create where Match Only belongs.** Documents that create clients or projects freely produce duplicates that structured data then cannot fix.
* **No folder filters.** Every file is scored against every type; accuracy drops and reviews pile up.
* **Metadata mode on a family that answers questions.** No entities or relationships are extracted in those modes.

## Exit criteria

* [ ] Every document family is an artifact type or explicitly out of scope (None)
* [ ] Every sample classifies to the right type at 0.8 or above in **Classify Only**
* [ ] **Classify & Extract** results reviewed with the SME for at least three samples per family
* [ ] Processing type agreed for every entity, and the Attach Document entity chosen (one per type)
* [ ] Extraction policy set and justified per family
* [ ] Folder filters list the allowed artifact types
* [ ] Artifact types show **valid** in **Model & Define > Compatibility**

## Next

[Workshop 6: Data Mapping](/implementation/ws-data-mapping) maps the structured exports that documents will match against.
