Two pipelines, one graph
Experio builds the graph from two kinds of source. They share the ontology but work differently, and most accuracy problems come from not planning for that difference.The document pipeline
1. Connect and scan
A connector authorises access to a storage provider (Box, Google Drive or SharePoint). A data source says which folders to scan and how. Its settings include recursion, file types, OCR, how many pages to read for classification, and folder filters. A folder filter limits which artifact types a file in that folder can be. For example, files in the Resumes library can only be Resumes. Starting a scan records every file found. The scan is refused if any artifact type it depends on is invalid against the current ontology (see ontology compatibility).2. Download and parse
Each file is downloaded and converted to text. Files with no extractable text are skipped. Enable OCR on the data source if the library contains scanned PDFs.3. Classify
An AI model scores the file against each allowed artifact type, and against “None”, using each type’s classification instructions.- A score of 0.8 or more assigns the type. This threshold is fixed.
- A lower score sends the file to a classification review in Process > Conflict Resolution. A reviewer picks the type or re-runs classification with extra guidance.
- “None”, or a type the folder filter doesn’t allow, means the file is skipped.
4. Extract
The model reads the text and extracts only the entity types, attributes and relationships listed on the artifact type. The artifact type’s list is a subset of the ontology. The prompt includes:- each attribute’s extraction instructions (including the expected format), taken from the ontology or overridden on the artifact type;
- the expanded values of any
@TaxonomyNamereferenced in those instructions, so values come from a controlled list; - the client firm’s own name and synonyms from Client Configuration, so the firm isn’t extracted as an outside company.
The policy also sets the model tier (large, medium or small) and whether the validation pass runs. Very large files are automatically downgraded to metadata only, to control cost.
5. Match
Each extracted entity is compared with what’s already in the graph. The artifact type says what to do with each entity type:
A relationship can be marked required upstream. The relationship’s source entity (the parent) is then resolved first, and the target is matched in its context. For example, with
Contract IMPOSES Obligation marked, the contract is resolved first, so a “monthly status report” obligation is matched only among that contract’s obligations, not every contract’s. The parent entity type needs a vector index for the option to be available.
The matching strategy for the entity type sets the methods (exact, fuzzy, phonetic, synonym, vector similarity), normalisation (strip “Inc”, “LLC”, punctuation) and thresholds. The defaults are:
6. Write to the graph
Nodes are written with their ontology label, and their attributes are converted to the ontology’s types (text, number, date, boolean, list, enum). Relationships are written with their attributes. Each file also becomes a Document node that holds its text and links to the entities found in it. Entity types with a vector index (Ontology > Indexes tab) get embeddings, which chat and matching use for similarity search. Every node and relationship gets lineage: which file, row, rule or reviewer created or changed it. Users see it through View lineage on citations.The structured pipeline
- Sources. CSV, Excel and JSON files from a connected provider or an upload, or a REST API with key, bearer, basic or OAuth2 authentication and pagination. There is no direct database connector. Database data arrives as an export file or through an API.
- Node mappings map a column to an entity type and attribute, with an operation: create, match or create-or-match. The match property (default
name) is the key used to find an existing node. - Relationship mappings name the relationship, the entity type at each end, and the column and property used to find each end.
- Incremental sync uses an optional record ID column and a timestamp column.
Where the two pipelines meet
A structured record and a document mention become one node when:- they use the same entity type, for example both are
Client; - the structured load uses a stable key as its match property, for example the CRM’s canonical account name, a project code or an email address;
- the document side is set to Create or Match or Match Only for that type, with a matching strategy that can bridge the variants (fuzzy, synonyms, suffix removal, vector similarity);
- key values are formatted the same way everywhere, for example
NB-2024-117, notNB2024117in one file andNB-2024-117in another.
After ingestion: enrichment
Enrichment rules add what the sources don’t state outright. Examples are tagging each project with an industry from its client and description, or inferring the skills a project used. They run over nodes already in the graph, on demand or as a step in a flow. They don’t run automatically after each file.Settings that affect accuracy
To check accuracy, use Process > Graph Evaluation for extraction quality and golden questions for answer quality. See Accuracy Validation.