At a glance
Before the workshop
FFE and Experio SMEs- From the data inventory, list the document families and where each lives. Propose a first grouping.
- Create a draft artifact type per family with a name and one-line classification instruction, so the workshop starts with something to test.
- Upload the sample documents to a File Upload data source and run a job, so they are parsed and available to the AI Assistant.
- Collect 3–5 real documents per family: a typical one, an old-format one, a short one, and one that is easy to confuse with another family (for example an SOW that contains full contract terms).
- Note how you tell the family apart from its look-alikes. “Our SOWs always reference an MSA and have a project code in the header” is exactly what the classifier needs.
Agenda
Running the workshop
1. Group the corpus into families
A family is a set of documents that serve the same purpose and answer the same questions. Group by purpose, not by file format or folder. Ask: “Which golden questions does this kind of document answer?” A family that answers no question does not need an artifact type; it will be classified as None and skipped.2. Write classification instructions
The classifier scores every allowed artifact type from 0 to 1 against its classification instructions. A document is assigned a type only if its best score reaches 0.8. This threshold is fixed. Below it, a classification review goes to Process > Conflict Resolution and the file waits for a person to choose. A result of None, or a type not allowed for that folder, means the file is skipped. Good instructions have three parts:- What it is: purpose and typical titles.
- How to recognise it: headings, clauses, layout, who signs it.
- How to tell it from look-alikes: name the confusable families and the deciding feature.
3. Choose what to extract
Pick the subset of the ontology this family can actually supply. Ask: “If this document were the only source, which golden-question facts would it give us?” Extract only those.- Select the entities and, per entity, the attributes. Extraction instructions default from the ontology; override them here when this family words things differently.
- Select the relationships between the chosen entities, with their attributes (for example
role). - Mark attributes required only if the document always has them.
4. Decide the processing type per entity
Load structured master data first (Workshop 6), then let documents Create or Match or Match Only against it. Matching detail is Workshop 7.
Mark a relationship Required upstream dependency when the child should resolve in the context of its parent. The relationship’s source is the parent, and the parent type needs a vector index (set in Workshop 3).
5. Set the extraction policy
The extraction policy is a cost and accuracy trade-off per family.
In the two metadata modes Experio skips entity extraction. It writes one shell node of the attached (or first) entity type, named after the file, plus the snippet when there is one, and no relationships. The parsed text is still stored on the Document node for chat. A cost guard also drops very large files to metadata-only. See Extraction Policy.
6. Map folders to allowed types
Each data source filter can list the artifact types allowed for the files it matches. The allowed set is the union across matching filters, or every type if no filter lists any. Narrowing the set is the cheapest accuracy gain there is: a file in the Resumes library cannot be misclassified as an MSA. Ask the IT admin: “Is each folder dedicated to one kind of document, or mixed?”Worked example: Northbridge Consulting
Everything else (templates, invoices, internal memos) is classified as None and skipped. The Resumes library filter allows Resume only; the Engagements filter allows Proposal, Statement of Work, Master Services Agreement, Case Study and Status Report.
Spec: Statement of Work
Helen Park brought four SOWs, one amendment and one MSA. The first draft scored the amendment 0.85 as a SOW and the MSA 0.7 as a SOW, so the room added look-alike rules.
Project is Match Only: a SOW must never create a project that finance does not know. Northbridge is kept out of Client automatically, because the firm name and its synonyms from Client Configuration are sent with every extraction.
Spec: Resume
Sam Whitfield’s samples came in three layouts. The classification instruction was short, because the Resumes folder allows nothing else:
The room decided not to extract projects from resumes. Resume project names (“Lakeshore cloud work”) cannot be matched to project codes, and
assignments.csv already holds the truth for who worked on what.
Accelerate with AI
Available today: Artifact Types > AI Assistant. Open an artifact type, open the AI Assistant, and choose up to 10 sample files that a job has already parsed. It drafts the extraction schema against the existing ontology, and you refine it by chatting (“drop Obligation, add liability_cap”). It needs an ontology first. Use it to produce the first draft before the workshop, then correct it with the SMEs using Test Ingestion.In development — availability depends on your release.
questions.md) and proposes artifact-type changes as cards, which a person reviews, applies and saves on the Artifact Types page. The manual path always works.
Entering it in Experio
1
Create the artifact type
Go to Model & Define > Artifact Types and create a type. Enter the name and classification instructions, and set Ingestion extraction (mode, model tier, validation pass, Excel mode).
2
Add entities and relationships
Add each entity with its processing type and, for one entity, Attach document to this entity. Choose attributes and override extraction instructions where needed. Add relationships, their attributes and any required upstream dependency.
3
Test
Click Test Ingestion, upload a sample, and run Classify Only first, then Classify & Extract. Review classification scores, entities, relationships and raw JSON. Use Max Pages to test long documents quickly.
4
Map folder filters
In Connect > Data Sources, edit each filter and choose the artifact types allowed for the files it matches.
Common pitfalls
- Instructions that only describe. Without “how to tell it from X”, look-alikes score close together and land in review.
- Testing only typical samples. Bring the awkward ones; they are what fails in production.
- Extracting everything. Each extra entity is cost and a chance of noise. Extract what the questions need.
- Create where Match Only belongs. Documents that create clients or projects freely produce duplicates that structured data then cannot fix.
- No folder filters. Every file is scored against every type; accuracy drops and reviews pile up.
- Metadata mode on a family that answers questions. No entities or relationships are extracted in those modes.
Exit criteria
- Every document family is an artifact type or explicitly out of scope (None)
- Every sample classifies to the right type at 0.8 or above in Classify Only
- Classify & Extract results reviewed with the SME for at least three samples per family
- Processing type agreed for every entity, and the Attach Document entity chosen (one per type)
- Extraction policy set and justified per family
- Folder filters list the allowed artifact types
- Artifact types show valid in Model & Define > Compatibility