At a glance
Pre-flight checklist
Do not start a job until every item is true. Each one prevents a class of failure that is tedious to clean up afterwards.Step 1: Load structured master data first
Structured loads create the backbone that documents attach to: clients, people, projects. Structured matching is exact-value only, and a relationship mapping never creates its endpoints. So the order matters. Load the parent entity types before anything that points at them.1
Run each structured data source in the agreed load order
From Process > Jobs, click Start New Job, pick the data source and choose Full Scan. Wait for each job to finish before you start the next.
2
Check that the rows became nodes
Compare the row count of each file with what landed. Ask the pilot assistant simple counting questions (“How many clients are in the graph?”) and compare the answer with the export.
3
Check that the relationships landed
Pick five records you know well and ask about their relationships (“Which projects are for Lakeshore Health?”). If a relationship is missing, the usual cause is a key that differs between two files, for example a trailing space or a different email case.
4
Spot-check lineage
From a cited entity in chat, open View lineage. Each structured record should show the data source and job that created it. See lineage.
Step 2: Ingest a small document sample
Pick a sample that covers every artifact type and the messy cases, not just the tidy ones. A good default is 3 client folders and 100–300 files. Choose at least one client with a long history, one with an unusual name (abbreviations, “Inc.”, a merger), and one where the SMEs know the answers to several golden questions.- In Connect > Data Sources, open the document data source and narrow the folder filters to the sample folders. Use Test Filters to confirm the file list.
- Confirm each filter allows only the artifact types that belong there. The allowed types for a file are the union of the types on its matching filters, or all types if no filter matches.
- Start a Full Scan from Process > Jobs.
Watch the job
Open the job from Process > Jobs. The processing phases show where each file is: Download, Parse, Classify, Ingest. The file list can be filtered by status.
To retry selected files, select them in the job’s file list and use Requeue to send them back to a phase (download, parse, classify or ingest).
Step 3: Work the Conflict Resolution queues
Open Process > Conflict Resolution. The pilot is the best time to learn what lands here, because every review is a hint about what to tune. See conflict resolution.- Classification reviews. The file waits until someone picks the correct artifact type. If the same kind of document keeps landing here, the classification instructions for that type are too vague. Sharpen them rather than approving reviews forever.
- Match reviews. An extracted entity scored between 0.5 and 0.7 against an existing node. Choose Merge (same thing), Confirm New (different thing) or Skip. If you merge the same pair of spellings over and over, add a synonym or a normalization rule to the matching strategy.
The super user should work these queues during the pilot with the FFE beside them. After go-live, the queues are theirs.
Step 4: Spot-check and evaluate
Spot-check with lineage. For ten or so documents, open an entity cited in chat and use View lineage. Confirm that the node came from the right document and that attributes match the source text. Run Graph Evaluation. From the completed job, click Run Evaluation, or go to Process > Graph Evaluation. Set samples per artifact type. The default is 3. Use 10–20 during the pilot so that every artifact type gets a meaningful sample. The AI judge compares the source text, the artifact type schema and what landed in the graph, then rates each sample good, acceptable or poor, and lists issues: missed entities, hallucinated entities, incorrect attributes, wrong relationship endpoints, schema violations. See graph evaluation.Graph Evaluation only reports. It does not fix anything, and it only covers full-mode document jobs. Artifact types on
metadata_and_snippet or metadata_only are not evaluated in depth.Step 5: Tune and reprocess
Change one thing at a time, reprocess only the files it affects, and compare.
Use Test Ingestion on the artifact type to try an instruction change against one or two problem files before you reprocess the batch.
Keep a simple tuning log: date, what changed, which files you reprocessed, the before and after evaluation results. The log becomes the super user’s reference after handover.
Cost awareness
The pilot is where you calibrate cost per document, before anyone multiplies it by the full corpus.- Extraction policy. Each artifact type has a mode (
full,metadata_and_snippet,metadata_only), a primary model tier (large, medium, small) and a validation pass toggle. Put high-volume, low-value types such as status reports onmetadata_and_snippet. See extraction policy and the Extraction Policy admin page. - Cost guard. Very large files are downgraded to metadata-only automatically when their estimated chunk count passes
INGESTION_COST_GUARD_CHUNK_THRESHOLD(default 120). If a big contract matters, check that it was not silently downgraded. - Model tiers. Start important types on the large tier. After evaluation is stable, try the medium tier on one type and re-run Graph Evaluation. Keep the downgrade only if quality holds.
- Excel. Exports and inventories rarely need full extraction. Use the Excel mode override or uncheck Ingest Excel files on export-only filters.
Worked example: Northbridge Consulting
The FFE and Marcus Lee ran the pilot over two weeks.- Pre-flight. Startup Health was green after one Ensure on a missing seed. Compatibility showed one stale matching strategy left over from the rename of Division to Practice in Workshop 3. Marcus reviewed it and clicked Mark as reviewed.
- Master data. They loaded
employees.csv,accounts.csv,projects.csvand thenassignments.csv.employees.csvhad 652 rows and 652 Employee nodes appeared.assignments.csvcreated only 1,890 of 2,140 WORKS_ON_PROJECT relationships. The cause was that 250 rows used uppercase emails. Sam’s team fixed the export and the rerun was complete. - Document sample. They picked three client folders in the Engagements library (Lakeshore Health, Meridian Bank and a state agency), plus 40 resumes: 240 files in total.
- First results. 31 classification reviews, mostly SOW amendments scoring 0.6–0.7 between SOW and MSA. There were 18 match reviews: “Lakeshore Health, Inc.” and “LSH” against Lakeshore Health, and “Bob Chen” against Robert Chen.
- Tuning. They added “An amendment that changes an existing SOW is a Statement of Work” to the SOW classification instructions, and added “LSH” as a Client synonym. Then they requeued the affected files from Classify. Classification reviews dropped to 4.
- Graph Evaluation. With 15 samples per type, MSAs rated poor on liability caps (values in words were missed). An extraction instruction on
liability_cap(“convert amounts written in words to a number in USD”) fixed it on the next run. - Cost. Status reports were 45% of files but low value, so they stayed on
metadata_and_snippet. Resumes moved to the medium tier with no loss in evaluation quality.
When to expand from sample to full corpus
Expand in steps (sample, then one full library, then everything). Move to the next step only when:- Graph Evaluation rates most samples in each artifact type good or acceptable, with no repeated issue pattern.
- Classification reviews are rare: a few per hundred files, not a steady stream.
- Match reviews are mostly one-off cases, not the same spellings over and over.
- Cost per document is known and fits the budget when multiplied by the corpus size.
- The first golden questions round can run on this data (see Accuracy Validation).
Automate recurring loads with Flows
Once the order is stable, capture it in a flow so the super user doesn’t have to remember it. In Process > Flows, chain the structured scans in load order, then the document scan, then the enrichment rules from Workshop 8. For example, run “Tag project industry” after the projects load. Run it on demand during the pilot, and add a schedule later. See Flows.Common pitfalls
- Documents before master data. Documents then create Clients and Employees that the structured load later duplicates, because structured matching is exact-value only.
- Sampling only clean files. The pilot passes and the full corpus fails. Include scans, amendments and odd names on purpose.
- Approving reviews instead of fixing the cause. Every repeated review is a tuning signal.
- Editing instructions without reprocessing. Old files keep old results, and the evaluation looks unchanged.
- Changing several things at once. You can’t tell which change helped.
- Ignoring stale items in Compatibility. They don’t block scans, but they often hide a mapping that points at the wrong attribute.
Exit criteria
- Pre-flight checklist complete and recorded
- Structured master data loaded in order. Node and relationship counts match the source exports.
- Document sample ingested, with every failed and skipped file explained
- Conflict Resolution queues cleared, and repeated patterns turned into instruction or matching fixes
- Graph Evaluation run for every full-mode artifact type, with results at an acceptable level
- Extraction policy and model tier chosen per artifact type, with cost per document estimated
- Tuning log handed to the super user
- Flow created for the recurring load order