Skip to main content
Accuracy validation answers the question the sponsor actually cares about: can people trust what Experio tells them? It runs two complementary checks. Graph Evaluation asks whether the documents were read correctly. Golden-question scoring asks whether the answers are right. A graph can be accurately extracted and still give wrong answers, and the reverse happens too. You need both. This phase also teaches the client team to diagnose a wrong answer, which is the skill they’ll use most after go-live.

At a glance

Two checks, two different questions

Check A: Extraction accuracy with Graph Evaluation

Run Graph Evaluation on the latest pilot job for every full-mode artifact type, using 10–20 samples per type. For each type, record the share of samples rated good, acceptable and poor, plus the top issue categories. What it measures:
  • Whether entities and relationships supported by the document and the artifact type schema are in the graph (missed)
  • Whether graph facts are supported by the text (hallucinated)
  • Whether attributes and relationship endpoints are correct
What it does not measure:
  • Whether the schema is the right one. If the ontology has no attribute for a fact, the judge doesn’t expect it.
  • Structured data loads, and artifact types on metadata_and_snippet or metadata_only
  • Whether chat can find and use the data
Treat a repeated issue category as a signal, not a one-off. Three incorrect_attribute issues on expiration_date in MSAs point at one extraction instruction to fix.

Check B: Answer accuracy with golden questions

Experio has no built-in question scoring, so you score golden questions by hand in a shared sheet. Plan two or three rounds, with a round of fixes and reprocessing between each.
1

Prepare the sheet

Include one row per question with: ID, use case, question text, expected answer, where the truth lives, and columns for each round (answer, citations, score, root cause, fix).
2

Run each question in chat

Use the pilot assistant, as a pilot user with normal access, not as an admin. Start a new conversation for each question so earlier turns don’t influence the answer.
3

Record the answer

Paste the answer text, or a short summary for long generative answers, into the sheet.
4

Check against the expected answer

Did it contain the right entities, values and dates? Did it leave anything out? Did it add anything wrong?
5

Check citations and lineage

Open the citations. Does each one support the claim next to it? For a key fact, open View lineage to confirm which document or row produced it.
6

Score it

Pass: correct and complete, with supporting citations. Partial: mostly right, but something is missing, or a citation is weak. Fail: wrong, empty, or unsupported.
7

Assign a root cause

For every partial or fail, walk the decision tree below and record one primary root cause.
Agree the scoring rules with the SMEs before round 1. For generative questions such as a past-performance summary, score the facts and citations, not the writing style.

Root-cause decision tree

Start with one question: is the data needed for the answer actually in the graph? Check with a direct question in chat (“Show me the SOWs for Lakeshore Health and their expiration dates”), follow citations to their entities, and use View lineage.

Root cause to fix

Fixes to classification, extraction and matching only change newly processed files. Reprocess the affected files before the next round, or the score won’t move. See Pilot Ingestion.

Success criteria

Agree these with the sponsor at kickoff and restate them before round 1. A sensible default:
  • 80% or more of golden questions score pass, and no fails on questions the sponsor has marked critical
  • Every pass has at least one citation that supports its key fact
  • Every remaining partial or fail has a known root cause and either a planned fix or an accepted limitation
  • Graph Evaluation shows no repeated poor pattern for any full-mode artifact type
Score = passes / total questions. Report partials separately; don’t count them as half-passes, which hides problems.

Worked example: Northbridge Consulting

Northbridge had 45 golden questions across three use cases. Marcus and the FFE scored round 1 together, and Dana, Sam and Helen confirmed the expected answers. Examples from round 1:
  • Q1 “Which cloud migration projects did we deliver for healthcare clients since 2023, and who led each?” scored partial. The projects were right, but the leads included everyone on the project. The data was in the graph, so the root cause was query logic. The fix was a Cypher instruction: “Engagement lead = WORKS_ON_PROJECT role ‘Engagement Lead’”.
  • Q4 “Who has Azure data platform experience and has worked in healthcare?” scored fail. Resumes created new Employee nodes (“Bob Chen”) instead of matching HR records (“Robert Chen”), so the skills never met the assignments. The root cause was matching. The fixes were to enable phonetic matching for Employee and clear the match-review queue, then reprocess the resumes.
  • Q9 “Which contracts have a limitation-of-liability cap below $1M?” scored fail. Graph Evaluation had already flagged incorrect_attribute on liability_cap, so the root cause was extraction. The fix was the liability_cap extraction instruction from the pilot, followed by requeuing the MSAs from Ingest.
  • Q8 scored pass, but only for admins. Pilot users in Sales saw no citations. Access control was in shadow mode, but diagnostics showed Contracts would be hidden from them in enforce mode. Helen confirmed that was intended, so the question was re-scoped to Legal users.

Toward automated regression

Once the golden list is stable, ask Experio’s engineering team to convert it into automated regression scenarios. Each scenario holds a question and strings the answer is expected to contain. They can then rerun it after upgrades or ontology changes instead of scoring by hand. Write expected answers now in a form that converts easily: short, specific values (“NB-2024-117”, “Priya Shah”, “2026-03-31”) rather than prose.

Common pitfalls

  • Scoring as an admin. Admins see everything, so access-control problems stay hidden until go-live.
  • Reusing one conversation. Earlier turns change the answer and make rounds hard to compare.
  • Fixing without reprocessing. The same failure shows up in round 2.
  • Treating Graph Evaluation as the accuracy score. It covers extraction only. The sponsor’s measure is the golden questions.
  • Recording several root causes per question. Pick the first break in the chain. Fixing it often fixes the rest.
  • Moving the goalposts. Adding many new questions between rounds makes the trend meaningless. Add them to the next round as a separate group.

Exit criteria

  • Graph Evaluation run for every full-mode artifact type on the latest data, with results recorded
  • At least two scoring rounds completed and recorded in the shared sheet
  • Success criteria met, or the gaps accepted in writing by the sponsor
  • Every partial or fail has a root cause and either a fix applied or an accepted limitation
  • Access control (if used) tested with three or more users at different levels
  • Golden list and expected answers updated and owned by the super user

Next

Continue to Go-Live & Handover.