> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Pilot Ingestion

> Load structured master data, then a small document sample, and tune the model until the graph looks right

Pilot ingestion is where the model you designed in the workshops meets real data for the first time. You load a small, deliberate slice of the client's content, look closely at what lands in the graph, and fix the configuration before anyone asks a question in chat. It is much cheaper to fix a classification instruction after 200 files than after 20,000.

The goal is not volume. The goal is a graph you trust for a representative sample, and a tuning loop the super user understands well enough to run without you.

## At a glance

| | |
| - | - |
| **Purpose** | Prove the model on real data: structured master data first, then a document sample. Tune until results are stable. |
| **When** | Phase 3, weeks 4–6. Workshop 8 (Enrichment Rules) runs in this phase once real nodes exist. |
| **Who** | FFE (drives), super user (reviews queues and spot-checks), SMEs (answer "is this right?"), client IT admin (connector access) |
| **Inputs** | Signed-off ontology, taxonomies, artifact types, data mappings and matching strategies from Workshops 3–7 |
| **Outputs** | Master data loaded, a sample of 100–300 documents ingested, empty or near-empty conflict queues, Graph Evaluation runs per artifact type, a tuning log |
| **Admin pages** | Observe > Startup Health, AI & Agents > Model Configurations, Model & Define > Compatibility, Administer > Client Configuration, Connect > Connectors, Connect > Data Sources, Process > Jobs, Process > Conflict Resolution, Process > Graph Evaluation, Process > Flows |

## Pre-flight checklist

Do not start a job until every item is true. Each one prevents a class of failure that is tedious to clean up afterwards.

| Check | Where | Why it matters |
| - | - | - |
| Startup Health is green: run **Verify all**; every check is **ready** | Observe > Startup Health | Missing seeds or lineage indexes make ingestion fail or skip lineage, and without lineage you can't evaluate the pilot. |
| Ingestion model tiers are set (large, plus small for secondary passes; medium if you plan to use it) | AI & Agents > Model Configurations, then Administer > System Settings | Extraction calls fail if a tier points at nothing. |
| Every dependent config is **valid** | Model & Define > Compatibility | An **invalid** artifact type, mapping or rule blocks the scan. **Stale** only warns, but you should review it anyway. See [ontology compatibility](/implementation/key-concepts#ontology-compatibility). |
| Client firm name and synonyms are entered | Administer > Client Configuration | Without this, the client's own name gets extracted as an external company on every document. See [client configuration](/implementation/key-concepts#client-configuration). |
| Connectors are authorised and data sources validate | Connect > Connectors, Connect > Data Sources | An expired OAuth token leaves files in **Authentication Failed** until someone re-authorises. |
| Folder filters limit the scan to the sample | Connect > Data Sources (Configure Filters, then Test Filters) | Stops you from ingesting the whole library by accident. |

<Warning>
  If Startup Health reports lineage indexes as **terminal**, schedule a maintenance window for **Force ensure** before you ingest. Do not run it while ingestion is writing to the graph.
</Warning>

## Step 1: Load structured master data first

Structured loads create the backbone that documents attach to: clients, people, projects. [Structured matching](/implementation/key-concepts#match-key) is exact-value only, and a relationship mapping never creates its endpoints. So the order matters. Load the parent entity types before anything that points at them.

<Steps>
  <Step title="Run each structured data source in the agreed load order">
    From **Process > Jobs**, click **Start New Job**, pick the data source and choose **Full Scan**. Wait for each job to finish before you start the next.
  </Step>

  <Step title="Check that the rows became nodes">
    Compare the row count of each file with what landed. Ask the pilot assistant simple counting questions ("How many clients are in the graph?") and compare the answer with the export.
  </Step>

  <Step title="Check that the relationships landed">
    Pick five records you know well and ask about their relationships ("Which projects are for Lakeshore Health?"). If a relationship is missing, the usual cause is a key that differs between two files, for example a trailing space or a different email case.
  </Step>

  <Step title="Spot-check lineage">
    From a cited entity in chat, open **View lineage**. Each structured record should show the data source and job that created it. See [lineage](/implementation/key-concepts#lineage).
  </Step>
</Steps>

<Tip>
  Keep a note of the exact counts you expect before you run the job (rows per file, distinct keys). That turns "looks about right" into a pass or fail check.
</Tip>

## Step 2: Ingest a small document sample

Pick a sample that covers every artifact type and the messy cases, not just the tidy ones. A good default is **3 client folders and 100–300 files**. Choose at least one client with a long history, one with an unusual name (abbreviations, "Inc.", a merger), and one where the SMEs know the answers to several golden questions.

1. In **Connect > Data Sources**, open the document data source and narrow the folder filters to the sample folders. Use **Test Filters** to confirm the file list.
2. Confirm each filter allows only the artifact types that belong there. The allowed types for a file are the union of the types on its matching filters, or all types if no filter matches.
3. Start a **Full Scan** from **Process > Jobs**.

### Watch the job

Open the job from **Process > Jobs**. The processing phases show where each file is: **Download**, **Parse**, **Classify**, **Ingest**. The file list can be filtered by status.

| Status you see | What it means | What to do |
| - | - | - |
| **Ingested** | Extracted and written to the graph | Spot-check a few |
| **Skipped** | Classified as "None" or as a type the folder filter does not allow, or excluded by an Excel rule | Check a handful. If a real proposal was skipped, fix the classification instructions or the filter. |
| **Pending Classification** that does not move | The top score was below the fixed 0.8 threshold, so a classification review is waiting | Resolve it in Conflict Resolution (below) |
| **Parse Failed** | The file could not be read (password protected, corrupt, image-only without OCR) | Enable **Use OCR** on the data source for scans, or exclude the file |
| **Download Failed** / **Authentication Failed** | Access problem | Ask the IT admin to check permissions or re-authorise the connector |
| **Classification Failed** / **Ingestion Failed** | A model call failed | Check **Observe > Logs**, then requeue the file |

To retry selected files, select them in the job's file list and use **Requeue** to send them back to a phase (download, parse, classify or ingest).

## Step 3: Work the Conflict Resolution queues

Open **Process > Conflict Resolution**. The pilot is the best time to learn what lands here, because every review is a hint about what to tune. See [conflict resolution](/implementation/key-concepts#conflict-resolution).

* **Classification reviews.** The file waits until someone picks the correct artifact type. If the same kind of document keeps landing here, the classification instructions for that type are too vague. Sharpen them rather than approving reviews forever.
* **Match reviews.** An extracted entity scored between 0.5 and 0.7 against an existing node. Choose **Merge** (same thing), **Confirm New** (different thing) or **Skip**. If you merge the same pair of spellings over and over, add a synonym or a normalization rule to the [matching strategy](/implementation/key-concepts#matching-strategy).

<Note>
  The super user should work these queues during the pilot with the FFE beside them. After go-live, the queues are theirs.
</Note>

## Step 4: Spot-check and evaluate

**Spot-check with lineage.** For ten or so documents, open an entity cited in chat and use **View lineage**. Confirm that the node came from the right document and that attributes match the source text.

**Run Graph Evaluation.** From the completed job, click **Run Evaluation**, or go to **Process > Graph Evaluation**. Set samples per artifact type. The default is 3. Use 10–20 during the pilot so that every artifact type gets a meaningful sample. The AI judge compares the source text, the artifact type schema and what landed in the graph, then rates each sample good, acceptable or poor, and lists issues: missed entities, hallucinated entities, incorrect attributes, wrong relationship endpoints, schema violations. See [graph evaluation](/implementation/key-concepts#graph-evaluation).

<Info>
  Graph Evaluation only reports. It does not fix anything, and it only covers full-mode document jobs. Artifact types on `metadata_and_snippet` or `metadata_only` are not evaluated in depth.
</Info>

## Step 5: Tune and reprocess

Change one thing at a time, reprocess only the files it affects, and compare.

| Symptom | Likely fix |
| - | - |
| Wrong artifact type, or many classification reviews | Sharpen the type's classification instructions. Add "is not" guidance. Tighten folder filters. |
| Missed entities or attributes | Improve attribute extraction instructions, including the expected format. Turn on the validation pass. |
| Hallucinated entities | Narrow the artifact type's entity list. Add "only extract if stated" to the instructions. |
| Duplicates (two "Lakeshore Health" nodes) | Add synonyms and suffix normalization to the matching strategy. Set Project to **Match Only**. |
| Wrong relationship endpoints | Mark the parent relationship **required upstream** so the parent resolves first |
| Firm extracted as an external company | Add the missing synonym in Client Configuration |

<Warning>
  Edits to classification or extraction instructions affect **future** processing only. Files already ingested keep their old results. To apply a change, reprocess the affected files: requeue them from the **Classify** or **Ingest** phase in the job's file list, or start a new Full Scan of the sample.
</Warning>

Use **Test Ingestion** on the artifact type to try an instruction change against one or two problem files before you reprocess the batch.

```mermaid theme={null}
flowchart LR
  A[Ingest sample] --> B[Watch Jobs:<br/>failed, skipped, waiting]
  B --> C[Work Conflict<br/>Resolution queues]
  C --> D[Spot-check lineage<br/>+ Graph Evaluation]
  D --> E{Issues<br/>acceptable?}
  E -- No --> F[Change one thing:<br/>instructions, filters,<br/>matching, policy]
  F --> G[Test Ingestion<br/>on problem files]
  G --> H[Requeue affected files]
  H --> B
  E -- Yes --> I[Expand sample or<br/>move to validation]
```

Keep a simple tuning log: date, what changed, which files you reprocessed, the before and after evaluation results. The log becomes the super user's reference after handover.

## Cost awareness

The pilot is where you calibrate cost per document, before anyone multiplies it by the full corpus.

* **Extraction policy.** Each [artifact type](/implementation/key-concepts#artifact-type) has a mode (`full`, `metadata_and_snippet`, `metadata_only`), a primary model tier (large, medium, small) and a validation pass toggle. Put high-volume, low-value types such as status reports on `metadata_and_snippet`. See [extraction policy](/implementation/key-concepts#extraction-policy) and the [Extraction Policy admin page](/admin-guide/extraction-policy).
* **Cost guard.** Very large files are downgraded to metadata-only automatically when their estimated chunk count passes `INGESTION_COST_GUARD_CHUNK_THRESHOLD` (default 120). If a big contract matters, check that it was not silently downgraded.
* **Model tiers.** Start important types on the large tier. After evaluation is stable, try the medium tier on one type and re-run Graph Evaluation. Keep the downgrade only if quality holds.
* **Excel.** Exports and inventories rarely need full extraction. Use the Excel mode override or uncheck **Ingest Excel files** on export-only filters.

## Worked example: Northbridge Consulting

The FFE and Marcus Lee ran the pilot over two weeks.

1. **Pre-flight.** Startup Health was green after one **Ensure** on a missing seed. Compatibility showed one stale matching strategy left over from the rename of Division to Practice in Workshop 3. Marcus reviewed it and clicked **Mark as reviewed**.
2. **Master data.** They loaded `employees.csv`, `accounts.csv`, `projects.csv` and then `assignments.csv`. `employees.csv` had 652 rows and 652 Employee nodes appeared. `assignments.csv` created only 1,890 of 2,140 WORKS\_ON\_PROJECT relationships. The cause was that 250 rows used uppercase emails. Sam's team fixed the export and the rerun was complete.
3. **Document sample.** They picked three client folders in the Engagements library (Lakeshore Health, Meridian Bank and a state agency), plus 40 resumes: 240 files in total.
4. **First results.** 31 classification reviews, mostly SOW amendments scoring 0.6–0.7 between SOW and MSA. There were 18 match reviews: "Lakeshore Health, Inc." and "LSH" against Lakeshore Health, and "Bob Chen" against Robert Chen.
5. **Tuning.** They added "An amendment that changes an existing SOW is a Statement of Work" to the SOW classification instructions, and added "LSH" as a Client synonym. Then they requeued the affected files from Classify. Classification reviews dropped to 4.
6. **Graph Evaluation.** With 15 samples per type, MSAs rated poor on liability caps (values in words were missed). An extraction instruction on `liability_cap` ("convert amounts written in words to a number in USD") fixed it on the next run.
7. **Cost.** Status reports were 45% of files but low value, so they stayed on `metadata_and_snippet`. Resumes moved to the medium tier with no loss in evaluation quality.

## When to expand from sample to full corpus

Expand in steps (sample, then one full library, then everything). Move to the next step only when:

* Graph Evaluation rates most samples in each artifact type **good** or **acceptable**, with no repeated issue pattern.
* Classification reviews are rare: a few per hundred files, not a steady stream.
* Match reviews are mostly one-off cases, not the same spellings over and over.
* Cost per document is known and fits the budget when multiplied by the corpus size.
* The first [golden questions](/implementation/key-concepts#golden-questions) round can run on this data (see [Accuracy Validation](/implementation/accuracy-validation)).

## Automate recurring loads with Flows

Once the order is stable, capture it in a [flow](/implementation/key-concepts#flow) so the super user doesn't have to remember it. In **Process > Flows**, chain the structured scans in load order, then the document scan, then the enrichment rules from Workshop 8. For example, run "Tag project industry" after the projects load. Run it on demand during the pilot, and add a schedule later. See [Flows](/admin-guide/flows).

## Common pitfalls

* **Documents before master data.** Documents then create Clients and Employees that the structured load later duplicates, because structured matching is exact-value only.
* **Sampling only clean files.** The pilot passes and the full corpus fails. Include scans, amendments and odd names on purpose.
* **Approving reviews instead of fixing the cause.** Every repeated review is a tuning signal.
* **Editing instructions without reprocessing.** Old files keep old results, and the evaluation looks unchanged.
* **Changing several things at once.** You can't tell which change helped.
* **Ignoring stale items in Compatibility.** They don't block scans, but they often hide a mapping that points at the wrong attribute.

## Exit criteria

* [ ] Pre-flight checklist complete and recorded
* [ ] Structured master data loaded in order. Node and relationship counts match the source exports.
* [ ] Document sample ingested, with every failed and skipped file explained
* [ ] Conflict Resolution queues cleared, and repeated patterns turned into instruction or matching fixes
* [ ] Graph Evaluation run for every full-mode artifact type, with results at an acceptable level
* [ ] Extraction policy and model tier chosen per artifact type, with cost per document estimated
* [ ] Tuning log handed to the super user
* [ ] Flow created for the recurring load order

## Next

Continue to [Accuracy Validation](/implementation/accuracy-validation) to score the golden questions against the pilot graph.
