Skip to main content
This page defines the terms used throughout the implementation. Each entry says what it is, what it’s used for, gives an example from Northbridge Consulting, and names where it lives in the admin.

How the pieces fit together

Modeling concepts

Knowledge graph

What it is. A database of nodes (things) connected by relationships (links). Each node has a type and attributes. What it’s for. Answering questions that cross sources and follow connections. For example: “who worked on healthcare projects for clients we have an MSA with”. Keyword search can’t answer that. Example. (Employee: Maria Gomez) -[WORKS_ON_PROJECT {role: "Engagement Lead"}]-> (Project: Lakeshore Cloud Migration) -[FOR_CLIENT]-> (Client: Lakeshore Health)

Ontology

What it is. The schema of the knowledge graph. It is the agreed list of the kinds of things the business cares about (entity types), the facts recorded about each (attributes), and the links allowed between them (relationships). What it’s for. Everything else is defined against it. Extraction only looks for what the ontology names. Data mappings only write to what it defines. Chat uses it to plan graph queries. A clear, question-driven ontology is the biggest single factor in answer accuracy. Example. Northbridge’s pilot ontology has eight entity types: Client, Project, Employee, Practice, Skill, Proposal, Contract and Obligation. It has thirteen relationships between them. Where. Model & Define > Ontology. Every save creates a revision (see Revision History). Built in Workshop 3.

Entity type

What it is. A kind of thing in the ontology, and the label on its nodes in the graph. It is written in singular PascalCase: Client, Project, Employee. What it’s for. It groups records that share attributes, relationships, a matching strategy and access rules. Rule of thumb. Make something an entity type when users ask questions about it, or when it has its own attributes and relationships. If it’s only a label on something else, make it an attribute, usually one with a taxonomy or an enum.

Attribute

What it is. A named, typed field on an entity type (or on a relationship). The types are text, number, date, boolean, list and enum (a fixed set of options). Each attribute can carry extraction instructions that tell the AI how to read it from documents, including the format you expect. What it’s for. Filters, sorts and facts in answers: dates, amounts, statuses, codes. Example. Contract.expiration_date (date; “the date the contract term ends, not the signature date”), Proposal.outcome (enum: won / lost / pending).

Relationship

What it is. An allowed, directed link between two entity types, named with an UPPER_SNAKE verb. Relationships can have their own attributes. What it’s for. Following connections: who worked on what, for whom, under which contract. Multi-hop questions depend entirely on relationships being modeled and populated. Example. Employee -[WORKS_ON_PROJECT {role, start_date, hours}]-> Project, Contract -[IMPOSES]-> Obligation.

Taxonomy

What it is. A controlled, usually hierarchical list of terms of one type. Examples are Industry, Service Line and Skills. Each taxonomy type holds items, which can nest (Healthcare > Providers) and be active or inactive. What it’s for. Consistency. When an attribute’s extraction instructions, or an enrichment rule, reference @Industry, the AI is given the taxonomy’s active leaf values and must pick from them. You get “Healthcare > Payers” every time instead of “health insurance”, “payer” and “HC payers”. Chat also sees taxonomies when it plans queries, so users can filter by them. Where it comes from. Often from another system: CRM picklists, an HR skills library, finance service-line codes, or industry standards such as NAICS. In that case it is imported as a CSV, not re-typed. Otherwise it’s designed in the workshop. Example. Northbridge’s Industry taxonomy is imported from the Salesforce Industry picklist. ObligationType is designed with Legal. Where. Model & Define > Taxonomies (CSV import on each taxonomy type’s page). Built in Workshop 4.
Taxonomies shape extraction and enrichment, not classification. They don’t decide which artifact type a document is.

Artifact type

What it is. A family of documents that are read the same way, for example Statement of Work, Resume or Master Services Agreement. It is also called a content type, in the Admin Guide and older screens. It holds:
  • classification instructions: how to recognise the family and tell it apart from look-alikes;
  • the subset of the ontology to extract: entity types, attributes and relationships, with instruction overrides;
  • per-entity processing type (Create Only / Create or Match / Match Only, plus Attach Document) and required upstream relationships (the parent is resolved first, so children match in its context);
  • an extraction policy.
What it’s for. It turns “read this PDF” into a precise, schema-bound task. A resume extracts the person and their skills. An SOW extracts the project, client, dates and value. Where. Model & Define > Artifact Types. Includes Test Ingestion for trying samples. Built in Workshop 5.

Extraction policy

What it is. A setting on each artifact type for how deep, and how expensive, extraction is. It sets the mode (full, metadata_and_snippet or metadata_only), the model tier (large, medium or small), the validation pass (on or off) and an Excel override. What it’s for. Spending AI effort where it answers questions. Contracts get full extraction on a large model. Thousands of status reports might get metadata and a snippet. Where. Model & Define > Artifact Types > Ingestion extraction. See Extraction Policy.

Data mapping

What it is. A column-by-column recipe that turns a structured source (a CSV, Excel or JSON file, or an API) into graph nodes, attributes and relationships. It uses node mappings (column → entity type and attribute, with create / match / create-or-match) and relationship mappings (which relationship, found by which columns at each end). What it’s for. Loading systems of record reliably, with no AI guesswork: the HR roster, CRM accounts, the finance project list, staffing assignments. Example. employees.csv maps work_email to Employee.email (match property) and manager_email to Employee -[REPORTS_TO]-> Employee. Where. Model & Define > Data Mapping (with Auto-map with AI). Built in Workshop 6.

Match key

What it is. The attribute used to decide that two records are the same thing: an ID, a code, an email or a canonical name. What it’s for. Joining sources. Structured loads match on it exactly. Documents match against the same nodes through their matching strategy. Choosing and cleaning match keys early prevents most duplicates. Example. Employee → email. Project → project_code. Client → canonical Salesforce account name.

Matching strategy

What it is. The rules for deciding whether an entity extracted from a document is the same as one already in the graph. It sets:
  • the methods and their weights: exact, fuzzy, phonetic, synonym and vector similarity;
  • normalisation: removing company suffixes, prefixes and punctuation;
  • filters: time window, related parent;
  • thresholds: auto-merge, AI decides, human review, new node.
There is a default strategy plus overrides per entity type. What it’s for. Preventing both duplicates (“Lakeshore Health” and “Lakeshore Health, Inc.” as two clients) and false merges (two different people named “J. Smith” combined). Where. Process > Matching Strategies. Built in Workshop 7.

Conflict resolution

What it is. The human review queues. Classification reviews hold files whose artifact type scored below 0.8. Match reviews hold entities whose match score fell in the human-review band, with Merge, Confirm New and Skip actions. What it’s for. Keeping a person in the loop exactly where the AI is unsure. The client super user usually owns these queues after go-live. Where. Process > Conflict Resolution.

Enrichment rule

What it is. An AI instruction that runs over nodes already in the graph and writes back an attribute, a new node, a relationship, or a node and a relationship. It can read the node, selected attributes, or its neighbours ({related.Client.industry}), and can be constrained to a taxonomy (@Industry). A relationship output links only to a target node that already exists; it doesn’t create the target. What it’s for. Business rules that are true but not written in any one source. For example, tag each project’s industry from its client, or infer the skills a project used from its scope. Where. Model & Define > Enrichment Rules. Runs on demand or as a step in a flow. Built in Workshop 8.

Client configuration

What it is. The firm’s own identity: its canonical name, synonyms and branding. What it’s for. Extraction is told who “we” are, so “Northbridge”, “NBC” and “Northbridge Consulting Group” aren’t extracted as an outside company or a client. Where. Administer > Client Configuration.

Ontology compatibility

What it is. A check of every artifact type, data mapping, enrichment rule and matching strategy against the current ontology revision. Each one is valid, stale (a rename was applied; please review), pending review or invalid (it points at something deleted). What it’s for. Safe change. Only invalid blocks scans and rules. Check it after every ontology save. Where. Model & Define > Compatibility. See Ontology Compatibility.

Pipeline concepts

Connector

What it is. An authorised connection to a storage provider (Box, Google Drive or SharePoint) or an API. Where. Connect > Connectors. Usually set up with the client’s IT admin.

Data source

What it is. What to read through a connector: the folders, recursion, file types, OCR, folder filters (which artifact types are allowed in which folders), and whether the source is documents or structured data. Where. Connect > Data Sources.

Flow

What it is. A scheduled or on-demand chain of data jobs: scan a data source, ingest, run enrichment rules, create structured jobs from folders. What it’s for. Keeping the graph current, for example a nightly scan followed by enrichment. It isn’t the same as an agent flow. Where. Process > Flows and Flow Executions.

Lineage

What it is. The recorded origin of every node and relationship: which document, table row, enrichment rule or reviewer created or changed it, and when. What it’s for. Trust and debugging. Users trace a citation back to its source. Implementers find out why a wrong fact is in the graph. Where. View lineage on citations and graph records. See Graph Lineage.

Graph evaluation

What it is. An AI judge that samples ingested documents per artifact type. It compares each document’s text with what landed in the graph and reports missed or invented entities and relationships, wrong attributes, wrong relationship endpoints and schema violations. What it’s for. Measuring extraction quality per artifact type. It reports only and changes nothing. It doesn’t test questions. Where. Process > Graph Evaluation. See Accuracy Validation.

Golden questions

What it is. The agreed list of real business questions, each with its expected answer and where the truth lives. It is gathered in Workshop 1 and kept as questions.md, a Markdown file in the engagement’s shared project folder. Start it from the template. What it’s for. It is the definition of success. It drives the ontology, and it is the test set for measuring answer accuracy before go-live and after every change. Scoring is a manual process today; see Accuracy Validation.

Consumption concepts

Assistant

What it is. A configured chat agent with a name, description, instructions, tools, model choices, and query instructions that encode business vocabulary. Where. AI & Agents > Agent Configuration, with shared rules in AI & Agents > AI Instructions. Designed in Workshop 10.

Agent flow

What it is. A multi-step AI workflow built on a visual canvas. Its blocks gather inputs, query the graph, run AI steps, pause for human review and produce files such as Word, PowerPoint, Excel or PDF. Users run it from chat. What it’s for. Repeatable jobs that produce an output, such as a past-performance write-up, a staffing shortlist or a contract renewal digest. Where. AI & Agents > Agent Flows. Reviews wait in Agent Inbox. See Agent Flows.

Document template

What it is. A branded Word, PowerPoint or Excel file, plus retrieval and output instructions, grouped by category (Case Study, Resume, Proposal and others). Where. Model & Define > Document Templates and Template Categories.

Persona

What it is. An optional audience profile, such as “Business development” or “New consultant”, that tailors intake questions and answers. It is turned on per deployment. Where. AI & Agents > Personas.

Access control

What it is. Record-level visibility rules on the graph: which entity types are private or public, how access rolls up (for example, a contract under its client), and who is granted access (for example, members of the project team). It runs in disabled, shadow or enforce mode. What it’s for. Ethical walls and confidential data. Answers and citations respect it. Where. Administer > Access Control. Designed in Workshop 9.

Knowledge-model copilot

In development. Availability depends on your release.
What it is. An Admin Copilot mode for the Ontology and Artifact Types pages. You attach sample documents and tables along with questions.md, and it proposes ontology and artifact-type changes. It can also draft a starting ontology from samples. Each proposal is a card that a person reviews, applies and saves. What it’s for. Speeding up Workshops 3 and 5. The workshop reviews a draft instead of starting from a blank page. The manual path is always available.