At a glance
Two checks, two different questions
Check A: Extraction accuracy with Graph Evaluation
Run Graph Evaluation on the latest pilot job for every full-mode artifact type, using 10–20 samples per type. For each type, record the share of samples rated good, acceptable and poor, plus the top issue categories. What it measures:- Whether entities and relationships supported by the document and the artifact type schema are in the graph (missed)
- Whether graph facts are supported by the text (hallucinated)
- Whether attributes and relationship endpoints are correct
- Whether the schema is the right one. If the ontology has no attribute for a fact, the judge doesn’t expect it.
- Structured data loads, and artifact types on
metadata_and_snippetormetadata_only - Whether chat can find and use the data
Check B: Answer accuracy with golden questions
Experio has no built-in question scoring, so you score golden questions by hand in a shared sheet. Plan two or three rounds, with a round of fixes and reprocessing between each.1
Prepare the sheet
Include one row per question with: ID, use case, question text, expected answer, where the truth lives, and columns for each round (answer, citations, score, root cause, fix).
2
Run each question in chat
Use the pilot assistant, as a pilot user with normal access, not as an admin. Start a new conversation for each question so earlier turns don’t influence the answer.
3
Record the answer
Paste the answer text, or a short summary for long generative answers, into the sheet.
4
Check against the expected answer
Did it contain the right entities, values and dates? Did it leave anything out? Did it add anything wrong?
5
Check citations and lineage
Open the citations. Does each one support the claim next to it? For a key fact, open View lineage to confirm which document or row produced it.
6
Score it
Pass: correct and complete, with supporting citations. Partial: mostly right, but something is missing, or a citation is weak. Fail: wrong, empty, or unsupported.
7
Assign a root cause
For every partial or fail, walk the decision tree below and record one primary root cause.
Agree the scoring rules with the SMEs before round 1. For generative questions such as a past-performance summary, score the facts and citations, not the writing style.
Root-cause decision tree
Start with one question: is the data needed for the answer actually in the graph? Check with a direct question in chat (“Show me the SOWs for Lakeshore Health and their expiration dates”), follow citations to their entities, and use View lineage.Root cause to fix
Success criteria
Agree these with the sponsor at kickoff and restate them before round 1. A sensible default:- 80% or more of golden questions score pass, and no fails on questions the sponsor has marked critical
- Every pass has at least one citation that supports its key fact
- Every remaining partial or fail has a known root cause and either a planned fix or an accepted limitation
- Graph Evaluation shows no repeated poor pattern for any full-mode artifact type
Worked example: Northbridge Consulting
Northbridge had 45 golden questions across three use cases. Marcus and the FFE scored round 1 together, and Dana, Sam and Helen confirmed the expected answers.
Examples from round 1:
- Q1 “Which cloud migration projects did we deliver for healthcare clients since 2023, and who led each?” scored partial. The projects were right, but the leads included everyone on the project. The data was in the graph, so the root cause was query logic. The fix was a Cypher instruction: “Engagement lead = WORKS_ON_PROJECT role ‘Engagement Lead’”.
- Q4 “Who has Azure data platform experience and has worked in healthcare?” scored fail. Resumes created new Employee nodes (“Bob Chen”) instead of matching HR records (“Robert Chen”), so the skills never met the assignments. The root cause was matching. The fixes were to enable phonetic matching for Employee and clear the match-review queue, then reprocess the resumes.
- Q9 “Which contracts have a limitation-of-liability cap below $1M?” scored fail. Graph Evaluation had already flagged
incorrect_attributeonliability_cap, so the root cause was extraction. The fix was theliability_capextraction instruction from the pilot, followed by requeuing the MSAs from Ingest. - Q8 scored pass, but only for admins. Pilot users in Sales saw no citations. Access control was in shadow mode, but diagnostics showed Contracts would be hidden from them in enforce mode. Helen confirmed that was intended, so the question was re-scoped to Legal users.
Toward automated regression
Once the golden list is stable, ask Experio’s engineering team to convert it into automated regression scenarios. Each scenario holds a question and strings the answer is expected to contain. They can then rerun it after upgrades or ontology changes instead of scoring by hand. Write expected answers now in a form that converts easily: short, specific values (“NB-2024-117”, “Priya Shah”, “2026-03-31”) rather than prose.Common pitfalls
- Scoring as an admin. Admins see everything, so access-control problems stay hidden until go-live.
- Reusing one conversation. Earlier turns change the answer and make rounds hard to compare.
- Fixing without reprocessing. The same failure shows up in round 2.
- Treating Graph Evaluation as the accuracy score. It covers extraction only. The sponsor’s measure is the golden questions.
- Recording several root causes per question. Pick the first break in the chain. Fixing it often fixes the rest.
- Moving the goalposts. Adding many new questions between rounds makes the trend meaningless. Add them to the next round as a separate group.
Exit criteria
- Graph Evaluation run for every full-mode artifact type on the latest data, with results recorded
- At least two scoring rounds completed and recorded in the shared sheet
- Success criteria met, or the gaps accepted in writing by the sponsor
- Every partial or fail has a root cause and either a fix applied or an accepted limitation
- Access control (if used) tested with three or more users at different levels
- Golden list and expected answers updated and owned by the super user