> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Accuracy Validation

> Check extraction quality with Graph Evaluation and answer quality with scored golden questions, then fix the root causes

Accuracy validation answers the question the sponsor actually cares about: *can people trust what Experio tells them?* It runs two complementary checks. **Graph Evaluation** asks whether the documents were read correctly. **Golden-question scoring** asks whether the answers are right. A graph can be accurately extracted and still give wrong answers, and the reverse happens too. You need both.

This phase also teaches the client team to diagnose a wrong answer, which is the skill they'll use most after go-live.

## At a glance

| | |
| - | - |
| **Purpose** | Measure extraction accuracy and answer accuracy, find the root cause of every failure, fix it, and re-score |
| **When** | Phase 4, weeks 6–7. Workshop 9 (Access Control) runs here if needed. |
| **Who** | FFE (runs scoring, diagnoses root causes), PM (tracks scores and fixes), super user (scores alongside), SMEs (judge the expected answers) |
| **Inputs** | Pilot graph from [Pilot Ingestion](/implementation/pilot-ingestion), the golden-question list from Workshop 1 with expected answers and where the truth lives, a pilot [assistant](/implementation/key-concepts#assistant) |
| **Outputs** | Scoring sheet per round, root-cause log, fixes applied, a final score against the agreed success criteria |
| **Admin pages** | Process > Graph Evaluation, Process > Jobs, Process > Conflict Resolution, AI & Agents > Agent Configuration, AI & Agents > AI Instructions, Administer > Access Control |

## Two checks, two different questions

| | Graph Evaluation | Golden-question scoring |
| - | - | - |
| **Question** | Did we extract what the document says? | Does the assistant answer the business question correctly? |
| **How** | An AI judge samples ingested documents per artifact type and compares source text, schema and graph | People run each question in chat and score the answer against the expected answer |
| **Built in?** | Yes: **Process > Graph Evaluation** | No. Scoring is manual today. |
| **Catches** | Missed or hallucinated entities and relationships, incorrect attributes, wrong endpoints, schema violations | Wrong or incomplete answers, bad citations, retrieval and query problems, missing data |
| **Misses** | Anything outside the schema, questions, chat retrieval, structured data, metadata-only types | Why the data is wrong (you diagnose that with lineage and evaluation) |

## Check A: Extraction accuracy with Graph Evaluation

Run [Graph Evaluation](/implementation/key-concepts#graph-evaluation) on the latest pilot job for every full-mode artifact type, using 10–20 samples per type. For each type, record the share of samples rated good, acceptable and poor, plus the top issue categories.

What it measures:

* Whether entities and relationships supported by the document and the artifact type schema are in the graph (**missed**)
* Whether graph facts are supported by the text (**hallucinated**)
* Whether attributes and relationship endpoints are correct

What it does **not** measure:

* Whether the schema is the right one. If the ontology has no attribute for a fact, the judge doesn't expect it.
* Structured data loads, and artifact types on `metadata_and_snippet` or `metadata_only`
* Whether chat can find and use the data

<Tip>
  Treat a repeated issue category as a signal, not a one-off. Three `incorrect_attribute` issues on `expiration_date` in MSAs point at one extraction instruction to fix.
</Tip>

## Check B: Answer accuracy with golden questions

Experio has no built-in question scoring, so you score [golden questions](/implementation/key-concepts#golden-questions) by hand in a shared sheet. Plan two or three rounds, with a round of fixes and reprocessing between each.

<Steps>
  <Step title="Prepare the sheet">
    Include one row per question with: ID, use case, question text, expected answer, where the truth lives, and columns for each round (answer, citations, score, root cause, fix).
  </Step>

  <Step title="Run each question in chat">
    Use the pilot assistant, as a pilot user with normal access, not as an admin. Start a new conversation for each question so earlier turns don't influence the answer.
  </Step>

  <Step title="Record the answer">
    Paste the answer text, or a short summary for long generative answers, into the sheet.
  </Step>

  <Step title="Check against the expected answer">
    Did it contain the right entities, values and dates? Did it leave anything out? Did it add anything wrong?
  </Step>

  <Step title="Check citations and lineage">
    Open the citations. Does each one support the claim next to it? For a key fact, open **View lineage** to confirm which document or row produced it.
  </Step>

  <Step title="Score it">
    **Pass**: correct and complete, with supporting citations. **Partial**: mostly right, but something is missing, or a citation is weak. **Fail**: wrong, empty, or unsupported.
  </Step>

  <Step title="Assign a root cause">
    For every partial or fail, walk the decision tree below and record one primary root cause.
  </Step>
</Steps>

<Note>
  Agree the scoring rules with the SMEs before round 1. For generative questions such as a past-performance summary, score the facts and citations, not the writing style.
</Note>

## Root-cause decision tree

Start with one question: *is the data needed for the answer actually in the graph?* Check with a direct question in chat ("Show me the SOWs for Lakeshore Health and their expiration dates"), follow citations to their entities, and use **View lineage**.

```mermaid theme={null}
flowchart TD
  A[Answer wrong or incomplete] --> B{Is the needed data<br/>in the graph?}
  B -- No --> C{Was the file or<br/>row ingested?}
  C -- No --> C1[Scope or access:<br/>filters, connector,<br/>failed or skipped file]
  C -- Yes --> D{Classified as the<br/>right artifact type?}
  D -- No --> D1[Classification instructions<br/>or folder filters]
  D -- Yes --> E{Were the entities and<br/>attributes extracted?}
  E -- No --> E1[Schema gap, extraction<br/>instructions, or<br/>extraction policy]
  E -- Yes --> F{Matched to the<br/>right node?}
  F -- No --> F1[Matching strategy:<br/>duplicates or wrong merge]
  F -- Yes --> G[Mapping or key mismatch<br/>for structured data]
  B -- Yes --> H{Did chat find the<br/>right entities?}
  H -- No --> H1[Entity resolution:<br/>vector index, synonyms,<br/>entity-resolution instructions]
  H -- Yes --> I{Was the query<br/>logic right?}
  I -- No --> I1[Cypher instructions:<br/>business definitions]
  I -- Yes --> J{Were citations<br/>missing for this user?}
  J -- Yes --> J1[Access control<br/>filtering citations]
  J -- No --> K[Assistant instructions:<br/>format, completeness]
```

### Root cause to fix

| Root cause | Typical evidence | Fix | Workshop to revisit |
| - | - | - | - |
| Not ingested (scope or access) | File absent from the job, or **Skipped** / **Download Failed** | Adjust folder filters, fix connector permissions, requeue the file | [WS2 Data Inventory](/implementation/ws-data-inventory) |
| Misclassified | Wrong artifact type in the job file list, or stuck in classification review | Sharpen classification instructions, tighten folder filters, reprocess | [WS5 Artifact Types](/implementation/ws-artifact-types) |
| Not extracted | Right type, but the entity or attribute is missing. Graph Evaluation reports `missed_*`. | Add the attribute to the ontology or artifact type, improve extraction instructions, check the extraction policy mode, reprocess | [WS3 Ontology](/implementation/ws-ontology), [WS5](/implementation/ws-artifact-types) |
| Wrongly matched or merged | Duplicate nodes, or facts from two clients on one node | Adjust the matching strategy (synonyms, normalization, thresholds). Use Match Only where a master list exists. | [WS7 Identity & Matching](/implementation/ws-identity-and-matching) |
| Not mapped | Structured rows missing, or relationships absent because the keys differ | Fix the data mapping or the export's key formatting, then reload | [WS6 Data Mapping](/implementation/ws-data-mapping) |
| Enrichment gap | Derived fact such as Project USED\_SKILL is absent | Run the rule, or refine its prompt | [WS8 Enrichment Rules](/implementation/ws-enrichment-rules) |
| Entity resolution in chat | Data exists, but the assistant didn't recognise "LSH" or "Bob Chen" | Enable a vector index on that entity type, add synonyms, add entity-resolution instructions | [WS7](/implementation/ws-identity-and-matching), [WS10](/implementation/ws-assistants-and-flows) |
| Query logic | Right entities, wrong filter ("led" read as any role) | Add Cypher instructions with the business definition | [WS10 Assistants & Agent Flows](/implementation/ws-assistants-and-flows) |
| Access control | Correct for an admin, but citations are missing for the pilot user | Review policies with **Preview as user**. Confirm the result is intended. | [WS9 Access Control](/implementation/ws-access-control) |
| Assistant instructions | Facts right, but the answer is truncated, badly formatted or hedged | Adjust assistant prompt extensions or AI Instructions | [WS10](/implementation/ws-assistants-and-flows) |
| Expected answer wrong | The SME's expected answer was out of date | Update the golden list. Don't count it against the system. | [WS1 Questions & Agents](/implementation/ws-questions-and-agents) |

<Warning>
  Fixes to classification, extraction and matching only change newly processed files. Reprocess the affected files before the next round, or the score won't move. See [Pilot Ingestion](/implementation/pilot-ingestion#step-5-tune-and-reprocess).
</Warning>

## Success criteria

Agree these with the sponsor at kickoff and restate them before round 1. A sensible default:

* **80% or more** of golden questions score **pass**, and **no fails** on questions the sponsor has marked critical
* Every pass has at least one citation that supports its key fact
* Every remaining partial or fail has a known root cause and either a planned fix or an accepted limitation
* Graph Evaluation shows no repeated **poor** pattern for any full-mode artifact type

Score = passes / total questions. Report partials separately; don't count them as half-passes, which hides problems.

## Worked example: Northbridge Consulting

Northbridge had 45 golden questions across three use cases. Marcus and the FFE scored round 1 together, and Dana, Sam and Helen confirmed the expected answers.

| Round | Pass | Partial | Fail | Pass rate | Main root causes |
| - | - | - | - | - | - |
| 1 | 28 | 9 | 8 | **62%** | Engagement lead not defined (Q1), resume names not matched (Q4, Q5), liability caps in words not extracted (Q9), SOW amendments misclassified (Q7) |
| 2 | 38 | 5 | 2 | **84%** | Remaining: two questions needing the opportunities export (not yet loaded), one generative summary missing a citation |

Examples from round 1:

* **Q1** "Which cloud migration projects did we deliver for healthcare clients since 2023, and who led each?" scored **partial**. The projects were right, but the leads included everyone on the project. The data was in the graph, so the root cause was query logic. The fix was a Cypher instruction: "Engagement lead = WORKS\_ON\_PROJECT role 'Engagement Lead'".
* **Q4** "Who has Azure data platform experience and has worked in healthcare?" scored **fail**. Resumes created new Employee nodes ("Bob Chen") instead of matching HR records ("Robert Chen"), so the skills never met the assignments. The root cause was matching. The fixes were to enable phonetic matching for Employee and clear the match-review queue, then reprocess the resumes.
* **Q9** "Which contracts have a limitation-of-liability cap below \$1M?" scored **fail**. Graph Evaluation had already flagged `incorrect_attribute` on `liability_cap`, so the root cause was extraction. The fix was the `liability_cap` extraction instruction from the pilot, followed by requeuing the MSAs from Ingest.
* **Q8** scored **pass**, but only for admins. Pilot users in Sales saw no citations. Access control was in shadow mode, but diagnostics showed Contracts would be hidden from them in enforce mode. Helen confirmed that was intended, so the question was re-scoped to Legal users.

## Toward automated regression

Once the golden list is stable, ask Experio's engineering team to convert it into automated regression scenarios. Each scenario holds a question and strings the answer is expected to contain. They can then rerun it after upgrades or ontology changes instead of scoring by hand. Write expected answers now in a form that converts easily: short, specific values ("NB-2024-117", "Priya Shah", "2026-03-31") rather than prose.

## Common pitfalls

* **Scoring as an admin.** Admins see everything, so access-control problems stay hidden until go-live.
* **Reusing one conversation.** Earlier turns change the answer and make rounds hard to compare.
* **Fixing without reprocessing.** The same failure shows up in round 2.
* **Treating Graph Evaluation as the accuracy score.** It covers extraction only. The sponsor's measure is the golden questions.
* **Recording several root causes per question.** Pick the first break in the chain. Fixing it often fixes the rest.
* **Moving the goalposts.** Adding many new questions between rounds makes the trend meaningless. Add them to the next round as a separate group.

## Exit criteria

* [ ] Graph Evaluation run for every full-mode artifact type on the latest data, with results recorded
* [ ] At least two scoring rounds completed and recorded in the shared sheet
* [ ] Success criteria met, or the gaps accepted in writing by the sponsor
* [ ] Every partial or fail has a root cause and either a fix applied or an accepted limitation
* [ ] Access control (if used) tested with three or more users at different levels
* [ ] Golden list and expected answers updated and owned by the super user

## Next

Continue to [Go-Live & Handover](/implementation/go-live).
