> ## Documentation Index
> Fetch the complete documentation index at: https://docs.experio.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Workshop 2: Data Inventory

> Find where the truth for each golden question lives, which keys link the sources, and what the pilot will load

This workshop decides which systems Experio will read, how the data will get out of them, and which shared keys will link the same client, person or project across those systems. Keys matter most. When the CRM, the project system and the document library each name a client differently and nothing links them, Experio builds three clients instead of one. Answers then come back incomplete, and the cause is hard to see. The keys you find here drive the ontology (WS3), the data mappings (WS6) and the matching rules (WS7).

## At a glance

| | |
| - | - |
| **Purpose** | Inventory every source behind the golden questions, find the linking keys, and choose the pilot sample |
| **When** | Phase 1 Discover, week 2, after all WS1 sessions |
| **Duration** | 2–3 hours |
| **Attendees** | Experio: FFE (facilitator), Experio SMEs as needed. Client: PM (scribe), super user, system owners (HR, Sales ops, Finance/PMO, Legal), IT admin for the document store |
| **Inputs** | Consolidated `questions.md` with a "where the truth lives" column; homework (below) |
| **Outputs** | Source inventory, key crosswalk, folder map, pilot sample, access and export requests with owners |
| **Admin pages** | Connect > Connectors, Connect > Data Sources (set up after the workshop) |

## Before the workshop

**FFE and Experio SMEs**

* Pull the "where the truth lives" column from `questions.md` into a first draft of the inventory table below. One row per system.
* Read the provider setup guides for the document store the client uses, so you can explain what IT must do (see [Connector prerequisites](#connector-prerequisites)).
* Prepare the key crosswalk table with the entity types the questions mention (client, person, project, contract).

**Client homework (send three working days before)**

Ask each system owner to bring:

1. A sample export of 20–50 real rows (CSV or XLSX), with the real column headers.
2. How the export is produced today (scheduled report, manual download, API) and how often.
3. Approximate row counts.
4. Known data quality problems ("inactive clients are never removed", "managers are stored by name, not ID").

Ask IT to bring:

1. A screenshot or listing of the top two or three levels of each document library.
2. Approximate file counts and file types per library.
3. Who can approve an OAuth app registration for Box, Google Drive or SharePoint.

## Agenda

| Time | Item | Lead |
| - | - | - |
| 0:00 | Recap: golden questions and the sources they named | FFE |
| 0:10 | Structured sources: walk each sample export | System owners |
| 0:50 | Unstructured sources: walk the library folder structure | IT admin, super user |
| 1:20 | Break | |
| 1:30 | Key crosswalk: how does each source identify the same thing? | FFE |
| 2:00 | Sensitive data, permissions and access prerequisites | IT admin |
| 2:20 | Pilot sample | Super user |
| 2:40 | Read-back and close | PM |

## Running the workshop

### 1. Map questions to sources

Go through `questions.md` and confirm the source for every must-have question. A question whose truth lives nowhere Experio can read cannot pass. Either find a source or downgrade it.

Ask the room:

* "When you answered this question last week, where did you look first? Where did you check?"
* "If two systems disagree, which one is right?" That system is the **system of record** for that fact.
* "Is there a spreadsheet someone keeps on the side that is more current than the system?"

### 2. Walk each structured source

Structured sources are systems with rows and columns: HR, CRM, finance and project systems, ERP. Experio reads them as [data sources](/implementation/key-concepts#data-source) with a structured type: CSV, XLSX/XLS or JSON files placed in Box, Google Drive, SharePoint or Dropbox, uploaded directly, or read from a REST API.

<Warning>
  Experio has **no direct database connector**. Data held in a database arrives as an export file or through an API. Agree now who produces each export, where it is dropped, and how often. A scheduled export to a fixed folder is the most reliable pattern.
</Warning>

For each source, ask:

* "What does one row represent? One employee? One assignment?"
* "Which column uniquely identifies a row? Does it ever change?"
* "Is there a column that shows when a row last changed?" (Useful for incremental sync; see [data mapping](/implementation/key-concepts#data-mapping).)
* "Which columns reference another system? A client name, a manager email?"
* "Which values are codes? Where is the list of valid codes?" (These are taxonomy candidates for WS4.)
* "Which rows should we ignore? Inactive, test, archived?"

### 3. Walk each document library

Unstructured sources are document libraries: proposals, contracts, resumes, reports. For each library, ask:

* "How is it organised? By client, by year, by document type?"
* "Is the structure consistent, or does every team do its own thing?"
* "Which folders hold final versions, and which hold drafts?"
* "Which folders should never be read?" (HR cases, personal folders, legal privilege.)
* "Are there scanned PDFs or images?" (OCR can be switched on per data source; it is slower.)

Folder structure matters because each data source can have **folder filters**, and a folder filter lists the [artifact types](/implementation/key-concepts#artifact-type) allowed in it. A file is only classified against the types allowed for its folder. A clean "Resumes" library can allow only Resume, which makes classification faster and more accurate. You will set these in WS5, but you need the folder map now.

### 4. Build the key crosswalk

This is the most important part of the session. For each entity type the questions need, write down how every source identifies it.

Ask the room:

* "In the CRM, how do you identify Lakeshore Health? In the project system? In the folder name? In the contract text?"
* "Is there one ID that appears in more than one system?"
* "Is the name written the same way everywhere? With 'Inc.'? With abbreviations?"
* "For people: which systems have an email address, and which only have a name?"

Then decide:

* The **[match key](/implementation/key-concepts#match-key)** for each entity type: the one value that will identify it in the graph.
* Which source is the **master** for that entity type. That source is loaded first.

Why this matters: structured loads match **on exact values only**. `Lakeshore Health` and `Lakeshore Health, Inc.` are two different clients to a data mapping. Documents can match more flexibly (fuzzy, synonyms, AI and human review), but only if the master records already exist. See [matching strategy](/implementation/key-concepts#matching-strategy).

<Tip>
  If two structured exports reference each other by name, ask the system owner whether the export can include the ID instead. A one-line change to a report now saves weeks of matching problems later.
</Tip>

### 5. Sensitive data and permissions

Ask the room:

* "Which sources contain personal data, compensation, health information or privileged legal material?"
* "Which columns should never be loaded?" Drop them from the export rather than relying on the mapping to skip them.
* "Who may see what today? Is that rule written down?" If record-level rules exist (for example, only Legal and the project team see contracts), plan [Workshop 9: Access Control](/implementation/ws-access-control).

### 6. Choose the pilot sample

Pilot ingestion (weeks 4–6) loads a representative slice, not everything. A good pilot sample:

* Covers every must-have golden question end to end.
* Includes **full structured exports**. Master data is small and must be complete, or document matching has nothing to match against.
* Includes a **limited set of document folders**: typically 3 clients' folders, chosen so that they contain every artifact type and at least one messy case.
* Includes the resumes of the people staffed on those clients' projects.

Record the sample as folder paths, so the data sources in the pilot can be configured exactly.

## Worked example: Northbridge Consulting

### Source inventory

| Source | System | Type | Owner | Export path | Volume | Frequency | Key |
| - | - | - | - | - | - | - | - |
| Employees | Workday `employees.csv` | Structured | HR | Scheduled report to SharePoint "Experio Exports" | \~650 rows | Nightly | `employee_id`, `work_email` |
| Clients | Salesforce `accounts.csv` | Structured | Sales ops | Scheduled report to the same folder | \~400 rows | Nightly | `account_id`, account name |
| Opportunities | Salesforce `opportunities.csv` | Structured | Sales ops | Scheduled report | \~2,500 rows | Nightly | opportunity name, `account_id` |
| Projects | Deltek `projects.csv` | Structured | Finance/PMO | Manual export, weekly | \~1,800 rows | Weekly | `project_code` (for example `NB-2024-117`) |
| Assignments | Deltek `assignments.csv` | Structured | Finance/PMO | Manual export, weekly | \~14,000 rows | Weekly | `project_code` + `employee_email` |
| Proposals, SOWs, MSAs, case studies | SharePoint "Engagements" library | Unstructured | BD / Legal | OAuth connector | \~22,000 files | Daily changes | Client folder names |
| Resumes | SharePoint "Resumes" library | Unstructured | HR | OAuth connector | \~700 files | Monthly | Name and email in the document |
| Status reports | SharePoint, per project folder | Unstructured | PMO | OAuth connector | \~9,000 files | Weekly | Project code in the title |

### Key crosswalk

| Entity | Salesforce | Deltek | Workday | SharePoint | Document text | Decision |
| - | - | - | - | - | - | - |
| Client | `account_id` 0015g00001, name "Lakeshore Health" | `client_name` "Lakeshore Health, Inc." | – | Folder `Engagements/Lakeshore Health` | "Lakeshore Health System", "LSH" | Match key: canonical Salesforce name. Deltek export changed to use the Salesforce name. Synonyms and suffix removal in WS7 |
| Project | Opportunity links to project (sometimes) | `project_code` NB-2024-117 | – | Subfolder `NB-2024-117 EHR Cloud Migration` | Code usually in SOW header and status report title | Match key: `project_code`. Documents may not create projects |
| Employee | – | `employee_email` | `employee_id`, `work_email` | Resume file name "Robert Chen CV.docx" | "Bob Chen" | Match key: email for structured loads; documents match by name with human review |
| Contract | – | – | – | Under client folder, `Contracts/` | Title, parties, effective date | Created from documents only |

Findings logged as decisions:

* **D-006**: Load order is employees, then accounts, then projects, then assignments, then documents.
* **D-007**: Finance adds the Salesforce account name as a column in `projects.csv` so Project FOR\_CLIENT Client can match exactly.
* **D-008**: `manager_email` added to `employees.csv` (it previously held the manager's name only).
* **D-009**: Compensation columns removed from the Workday export at source.

### Pilot sample

* Full exports: `employees.csv`, `accounts.csv`, `projects.csv`, `assignments.csv`.
* Documents: `Engagements/Lakeshore Health`, `Engagements/Meridian Bank`, `Engagements/State of Illinois DHS` (together they contain every artifact type).
* Resumes of the \~120 people assigned to those clients' projects.
* `opportunities.csv` deferred to after the pilot (optional mapping).

## Connector prerequisites

Document libraries are read through a [connector](/implementation/key-concepts#connector): an authorised connection to Box, Google Drive or SharePoint that uses OAuth. The client's IT admin must register an app with the provider before the FFE can create the connector. Start these requests at the end of this workshop. Approvals often take a week or more.

<Tabs>
  <Tab title="SharePoint">
    Register an app in Microsoft Entra ID (Azure AD), single tenant, with a **Web** redirect URI that Experio provides (ending `/admin/datasource/sharepoint/oauth_callback`). Add Microsoft Graph **delegated** permissions `Sites.Read.All`, `Files.Read.All`, `User.Read` and `offline_access`, then grant admin consent. Provide the client ID, client secret and tenant ID. Requires a SharePoint or Global Administrator. See the SharePoint setup guide in `docs/providers/SharePoint-Setup.md`.
  </Tab>

  <Tab title="Box">
    Create a Custom App with **User Authentication (OAuth 2.0)** in the Box Developer Console, add the exact redirect URI Experio provides (ending `/admin/datasource/box/oauth_callback/`), select the scopes in the Box guide, and provide the client ID and secret. Enterprise policy may require a Box admin to approve the app. See `docs/providers/box-app-configuration.md`.
  </Tab>

  <Tab title="Google Drive">
    In Google Cloud Console, enable the Google Drive API, configure the OAuth consent screen with `drive.readonly`, and create a **Web application** OAuth client with the redirect URI Experio provides (ending `/admin/datasource/google-drive/oauth_callback/`). A service-account option also exists for Shared Drives. See `docs/providers/google-drive-app-configuration.md`.
  </Tab>
</Tabs>

The account used to authorise the connector must be able to read every folder in the pilot sample. Agree with IT which account that is.

## Entering it in Experio

Most outputs go into the decision log and the inventory. Connectors and data sources are usually created after WS3–WS6, once the model is agreed, but you can create and test the connector early to prove access.

1. **Connect > Connectors**: create the connector with the credentials IT provided, then **Authorize** and **Test**. See [Connectors](/admin-guide/connectors).
2. **Connect > Data Sources**: create one data source per library, with the pilot folder paths. Use **Test Filters** to check that the right files match. See [Data Sources](/admin-guide/data-sources).
3. Structured exports become data sources with a structured type once their mappings exist (WS6).

## Common pitfalls

* **Assuming a database connection.** There is none. Agree export files or an API in this session.
* **Names as keys.** If the only link between two exports is a free-text name, fix the export before the pilot.
* **Sampling documents without master data.** Documents matched against an empty graph create duplicate clients and people.
* **Forgetting drafts.** Libraries full of `v3 FINAL (2).docx` produce conflicting facts. Point data sources at final-version folders where possible.
* **Late OAuth approvals.** The app registration is often the longest lead-time item in the whole implementation.
* **Loading sensitive columns "just in case".** Remove them at source.

## Exit criteria

* [ ] Every must-have golden question has a confirmed source
* [ ] Inventory complete: owner, format, export path, volume, frequency, access, sensitive data, quality issues
* [ ] Key crosswalk complete, with a match key and master source per entity type
* [ ] Export changes agreed with owners and dates
* [ ] Folder map recorded for each document library
* [ ] Pilot sample agreed as folder paths and export files
* [ ] OAuth app registration requested, with an IT owner and date
* [ ] Need for WS9 Access Control decided

## Next

Continue to [Workshop 3: Ontology](/implementation/ws-ontology).
