At a glance
Before the workshop
FFE and Experio SMEs- Pull the “where the truth lives” column from
questions.mdinto a first draft of the inventory table below. One row per system. - Read the provider setup guides for the document store the client uses, so you can explain what IT must do (see Connector prerequisites).
- Prepare the key crosswalk table with the entity types the questions mention (client, person, project, contract).
- A sample export of 20–50 real rows (CSV or XLSX), with the real column headers.
- How the export is produced today (scheduled report, manual download, API) and how often.
- Approximate row counts.
- Known data quality problems (“inactive clients are never removed”, “managers are stored by name, not ID”).
- A screenshot or listing of the top two or three levels of each document library.
- Approximate file counts and file types per library.
- Who can approve an OAuth app registration for Box, Google Drive or SharePoint.
Agenda
Running the workshop
1. Map questions to sources
Go throughquestions.md and confirm the source for every must-have question. A question whose truth lives nowhere Experio can read cannot pass. Either find a source or downgrade it.
Ask the room:
- “When you answered this question last week, where did you look first? Where did you check?”
- “If two systems disagree, which one is right?” That system is the system of record for that fact.
- “Is there a spreadsheet someone keeps on the side that is more current than the system?“
2. Walk each structured source
Structured sources are systems with rows and columns: HR, CRM, finance and project systems, ERP. Experio reads them as data sources with a structured type: CSV, XLSX/XLS or JSON files placed in Box, Google Drive, SharePoint or Dropbox, uploaded directly, or read from a REST API. For each source, ask:- “What does one row represent? One employee? One assignment?”
- “Which column uniquely identifies a row? Does it ever change?”
- “Is there a column that shows when a row last changed?” (Useful for incremental sync; see data mapping.)
- “Which columns reference another system? A client name, a manager email?”
- “Which values are codes? Where is the list of valid codes?” (These are taxonomy candidates for WS4.)
- “Which rows should we ignore? Inactive, test, archived?“
3. Walk each document library
Unstructured sources are document libraries: proposals, contracts, resumes, reports. For each library, ask:- “How is it organised? By client, by year, by document type?”
- “Is the structure consistent, or does every team do its own thing?”
- “Which folders hold final versions, and which hold drafts?”
- “Which folders should never be read?” (HR cases, personal folders, legal privilege.)
- “Are there scanned PDFs or images?” (OCR can be switched on per data source; it is slower.)
4. Build the key crosswalk
This is the most important part of the session. For each entity type the questions need, write down how every source identifies it. Ask the room:- “In the CRM, how do you identify Lakeshore Health? In the project system? In the folder name? In the contract text?”
- “Is there one ID that appears in more than one system?”
- “Is the name written the same way everywhere? With ‘Inc.’? With abbreviations?”
- “For people: which systems have an email address, and which only have a name?”
- The match key for each entity type: the one value that will identify it in the graph.
- Which source is the master for that entity type. That source is loaded first.
Lakeshore Health and Lakeshore Health, Inc. are two different clients to a data mapping. Documents can match more flexibly (fuzzy, synonyms, AI and human review), but only if the master records already exist. See matching strategy.
5. Sensitive data and permissions
Ask the room:- “Which sources contain personal data, compensation, health information or privileged legal material?”
- “Which columns should never be loaded?” Drop them from the export rather than relying on the mapping to skip them.
- “Who may see what today? Is that rule written down?” If record-level rules exist (for example, only Legal and the project team see contracts), plan Workshop 9: Access Control.
6. Choose the pilot sample
Pilot ingestion (weeks 4–6) loads a representative slice, not everything. A good pilot sample:- Covers every must-have golden question end to end.
- Includes full structured exports. Master data is small and must be complete, or document matching has nothing to match against.
- Includes a limited set of document folders: typically 3 clients’ folders, chosen so that they contain every artifact type and at least one messy case.
- Includes the resumes of the people staffed on those clients’ projects.
Worked example: Northbridge Consulting
Source inventory
Key crosswalk
Findings logged as decisions:
- D-006: Load order is employees, then accounts, then projects, then assignments, then documents.
- D-007: Finance adds the Salesforce account name as a column in
projects.csvso Project FOR_CLIENT Client can match exactly. - D-008:
manager_emailadded toemployees.csv(it previously held the manager’s name only). - D-009: Compensation columns removed from the Workday export at source.
Pilot sample
- Full exports:
employees.csv,accounts.csv,projects.csv,assignments.csv. - Documents:
Engagements/Lakeshore Health,Engagements/Meridian Bank,Engagements/State of Illinois DHS(together they contain every artifact type). - Resumes of the ~120 people assigned to those clients’ projects.
opportunities.csvdeferred to after the pilot (optional mapping).
Connector prerequisites
Document libraries are read through a connector: an authorised connection to Box, Google Drive or SharePoint that uses OAuth. The client’s IT admin must register an app with the provider before the FFE can create the connector. Start these requests at the end of this workshop. Approvals often take a week or more.- Box
- Google Drive
Entering it in Experio
Most outputs go into the decision log and the inventory. Connectors and data sources are usually created after WS3–WS6, once the model is agreed, but you can create and test the connector early to prove access.- Connect > Connectors: create the connector with the credentials IT provided, then Authorize and Test. See Connectors.
- Connect > Data Sources: create one data source per library, with the pilot folder paths. Use Test Filters to check that the right files match. See Data Sources.
- Structured exports become data sources with a structured type once their mappings exist (WS6).
Common pitfalls
- Assuming a database connection. There is none. Agree export files or an API in this session.
- Names as keys. If the only link between two exports is a free-text name, fix the export before the pilot.
- Sampling documents without master data. Documents matched against an empty graph create duplicate clients and people.
- Forgetting drafts. Libraries full of
v3 FINAL (2).docxproduce conflicting facts. Point data sources at final-version folders where possible. - Late OAuth approvals. The app registration is often the longest lead-time item in the whole implementation.
- Loading sensitive columns “just in case”. Remove them at source.
Exit criteria
- Every must-have golden question has a confirmed source
- Inventory complete: owner, format, export path, volume, frequency, access, sensitive data, quality issues
- Key crosswalk complete, with a match key and master source per entity type
- Export changes agreed with owners and dates
- Folder map recorded for each document library
- Pilot sample agreed as folder paths and export files
- OAuth app registration requested, with an IT owner and date
- Need for WS9 Access Control decided