> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# How Data Gets to the Graph

> The document and structured pipelines step by step, where they meet, and every setting that affects accuracy

## Two pipelines, one graph

Experio builds the graph from two kinds of source. They share the ontology but work differently, and most accuracy problems come from not planning for that difference.

| | Documents (unstructured) | Tables (structured) |
| - | - | - |
| **Examples** | Proposals, SOWs, contracts, resumes, status reports, case studies | HR, CRM and finance exports; JSON files; REST APIs |
| **Configured with** | [Artifact types](/implementation/key-concepts#artifact-type) | [Data mappings](/implementation/key-concepts#data-mapping) |
| **Who decides what becomes a node** | An AI model, limited to the artifact type's schema | The mapping, column by column |
| **How records are matched** | [Matching strategy](/implementation/key-concepts#matching-strategy): exact, fuzzy, phonetic, synonym and vector similarity, with AI and human review | Exact value of one match property, such as `email` or `project_code` |
| **Human review** | Classification and match reviews in **Process > Conflict Resolution** | None; failed rows are listed on the job |
| **Best at** | Detail, context, text: scope, obligations, experience | Reliable master lists and hard facts: who, which client, dates, amounts |

<Tip>
  **Load tables first, then documents.** Structured master data (people, clients, projects) creates clean records with reliable keys. When documents are ingested afterwards, their mentions of "Lakeshore Health" or "Bob Chen" match those records instead of creating near-duplicates.
</Tip>

## The document pipeline

```mermaid theme={null}
flowchart TD
    A["Connector + Data Source<br/>which provider, which folders"] --> B["Scan<br/>a record per file"]
    B --> C["Download"]
    C --> D["Parse to text<br/>OCR optional"]
    D --> E{"Classify<br/>which artifact type?"}
    E -->|"score 0.8 or more"| F["Extract<br/>entities, attributes and relationships<br/>allowed by the artifact type"]
    E -->|"score below 0.8"| R1["Classification review<br/>Conflict Resolution"]
    R1 -->|"reviewer picks a type"| F
    E -->|"None / not allowed"| S["Skipped"]
    F --> G{"Match each entity<br/>to existing records"}
    G -->|"0.9 or more"| M["Merge"]
    G -->|"0.7 – 0.9"| AI["AI decides"]
    G -->|"0.5 – 0.7"| R2["Match review<br/>Conflict Resolution"]
    G -->|"below 0.5"| N["New node"]
    M --> W[("Write to graph<br/>+ lineage")]
    AI --> W
    R2 --> W
    N --> W
```

### 1. Connect and scan

A [connector](/implementation/key-concepts#connector) authorises access to a storage provider (Box, Google Drive or SharePoint). A [data source](/implementation/key-concepts#data-source) says which folders to scan and how. Its settings include recursion, file types, OCR, how many pages to read for classification, and **folder filters**. A folder filter limits which artifact types a file in that folder can be. For example, files in the *Resumes* library can only be Resumes.

Starting a scan records every file found. The scan is refused if any artifact type it depends on is **invalid** against the current ontology (see [ontology compatibility](/implementation/key-concepts#ontology-compatibility)).

### 2. Download and parse

Each file is downloaded and converted to text. Files with no extractable text are skipped. Enable OCR on the data source if the library contains scanned PDFs.

### 3. Classify

An AI model scores the file against each allowed artifact type, and against "None", using each type's **classification instructions**.

* A score of **0.8 or more** assigns the type. This threshold is fixed.
* A lower score sends the file to a **classification review** in **Process > Conflict Resolution**. A reviewer picks the type or re-runs classification with extra guidance.
* "None", or a type the folder filter doesn't allow, means the file is skipped.

Accuracy here depends on well-written classification instructions, especially for look-alike documents such as SOWs and amendments. Folder filters that narrow the choice help too.

### 4. Extract

The model reads the text and extracts **only** the entity types, attributes and relationships listed on the artifact type. The artifact type's list is a subset of the ontology. The prompt includes:

* each attribute's **extraction instructions** (including the expected format), taken from the ontology or overridden on the artifact type;
* the expanded values of any `@TaxonomyName` referenced in those instructions, so values come from a controlled list;
* the client firm's own name and synonyms from **Client Configuration**, so the firm isn't extracted as an outside company.

Extraction runs as a primary pass, then an optional **validation pass** that looks for anything missed, then a relationship backfill for orphaned entities. Long documents are processed in chunks, and spreadsheets sheet by sheet.

How deep extraction goes is set by the [extraction policy](/implementation/key-concepts#extraction-policy) on each artifact type:

| Mode | What happens | Use for |
| - | - | - |
| `full` | AI extraction of the full schema | Documents that answer questions: SOWs, contracts, resumes |
| `metadata_and_snippet` | A lightweight record with a text snippet, no AI extraction | High-volume, low-value files you still want to find |
| `metadata_only` | A lightweight record only | Large exports, archives |

The policy also sets the model tier (large, medium or small) and whether the validation pass runs. Very large files are automatically downgraded to metadata only, to control cost.

### 5. Match

Each extracted entity is compared with what's already in the graph. The artifact type says what to do with each entity type:

| Processing type | Behaviour | Use for |
| - | - | - |
| **Create Only** | Always create a new node | Things that are unique to this document, such as an obligation clause |
| **Create or Match** | Run matching; merge if it's the same thing, otherwise create | Clients, people and skills mentioned in documents |
| **Match Only** | Update an existing node if found; never create | Things that must come from a system of record, such as projects that finance knows about |
| **Attach Document** (option) | Link the source Document to the entity | Resumes to the Employee, SOWs to the Project |

A relationship can be marked **required upstream**. The relationship's source entity (the parent) is then resolved first, and the target is matched in its context. For example, with `Contract IMPOSES Obligation` marked, the contract is resolved first, so a "monthly status report" obligation is matched only among that contract's obligations, not every contract's. The parent entity type needs a vector index for the option to be available.

The matching strategy for the entity type sets the methods (exact, fuzzy, phonetic, synonym, vector similarity), normalisation (strip "Inc", "LLC", punctuation) and thresholds. The defaults are:

| Score | Outcome |
| - | - |
| 0.9 and above | Merge automatically |
| 0.7 – 0.9 | An AI model decides |
| 0.5 – 0.7 | A reviewer decides in **Conflict Resolution** (Merge / Confirm New / Skip) |
| Below 0.5 | Create a new node (unless the entity type is Match Only) |

### 6. Write to the graph

Nodes are written with their ontology label, and their attributes are converted to the ontology's types (text, number, date, boolean, list, enum). Relationships are written with their attributes. Each file also becomes a **Document** node that holds its text and links to the entities found in it. Entity types with a vector index (Ontology > **Indexes** tab) get embeddings, which chat and matching use for similarity search.

Every node and relationship gets **lineage**: which file, row, rule or reviewer created or changed it. Users see it through **View lineage** on citations.

## The structured pipeline

```mermaid theme={null}
flowchart TD
    A["Structured data source<br/>CSV / Excel / JSON file or REST API"] --> B["Data mapping<br/>columns → entities, attributes, relationships"]
    B --> C{"Node mappings<br/>find by exact match property"}
    C -->|"found"| U["Update attributes<br/>+ lineage"]
    C -->|"not found"| K["Create node<br/>+ lineage"]
    U --> RM["Relationship mappings<br/>find both ends by exact value,<br/>then connect"]
    K --> RM
    RM --> G[("Knowledge graph")]
```

* **Sources.** CSV, Excel and JSON files from a connected provider or an upload, or a REST API with key, bearer, basic or OAuth2 authentication and pagination. There is no direct database connector. Database data arrives as an export file or through an API.
* **Node mappings** map a column to an entity type and attribute, with an operation: *create*, *match* or *create-or-match*. The **match property** (default `name`) is the key used to find an existing node.
* **Relationship mappings** name the relationship, the entity type at each end, and the column and property used to find each end.
* **Incremental sync** uses an optional record ID column and a timestamp column.

<Warning>
  Two rules shape every structured load:

  1. **Matching is exact.** `Lakeshore Health` and `Lakeshore Health, Inc.` are two different clients to a data mapping. Clean or standardise key columns before loading.
  2. **A relationship mapping never creates its endpoints.** Both ends must already exist, created by a node mapping in this load or an earlier one. Load in dependency order, for example people → clients → projects → assignments.
</Warning>

## Where the two pipelines meet

A structured record and a document mention become one node when:

1. they use the **same entity type**, for example both are `Client`;
2. the structured load uses a **stable key** as its match property, for example the CRM's canonical account name, a project code or an email address;
3. the document side is set to **Create or Match** or **Match Only** for that type, with a matching strategy that can bridge the variants (fuzzy, synonyms, suffix removal, vector similarity);
4. key values are **formatted the same way** everywhere, for example `NB-2024-117`, not `NB2024117` in one file and `NB-2024-117` in another.

The [Data Inventory](/implementation/ws-data-inventory) workshop finds the keys. The [Data Mapping](/implementation/ws-data-mapping) and [Identity & Matching](/implementation/ws-identity-and-matching) workshops set them up.

## After ingestion: enrichment

[Enrichment rules](/implementation/key-concepts#enrichment-rule) add what the sources don't state outright. Examples are tagging each project with an industry from its client and description, or inferring the skills a project used. They run over nodes already in the graph, on demand or as a step in a [flow](/implementation/key-concepts#flow). They don't run automatically after each file.

## Settings that affect accuracy

| Stage | Setting | Where |
| - | - | - |
| Classify | Classification instructions; folder filters; pages read for classification | Artifact Types; Data Sources |
| Extract | Which entity types, attributes and relationships each artifact type extracts; extraction instructions; `@Taxonomy` references | Ontology; Artifact Types; Taxonomies |
| Extract | Extraction mode, model tier, validation pass | Artifact Types > Ingestion extraction |
| Extract | Firm name and synonyms | Client Configuration |
| Match | Processing type per entity; required-upstream relationships | Artifact Types |
| Match | Methods, weights, normalisation, thresholds per entity type | Matching Strategies |
| Match | Vector index per entity type | Ontology > Indexes tab |
| Structured | Match property, operations, load order, clean keys | Data Mapping |
| Enrich | Rule instructions, inputs, outputs | Enrichment Rules |
| Answer | Assistant instructions, query and entity-resolution instructions | Agent Configuration; AI Instructions |
| Models | Which model runs each stage | Model Configurations |

To check accuracy, use **Process > Graph Evaluation** for extraction quality and [golden questions](/implementation/key-concepts#golden-questions) for answer quality. See [Accuracy Validation](/implementation/accuracy-validation).
